Video Keyframe Extraction Pipeline
Extract representative keyframes from videos with scene detection. Generate thumbnails, visual summaries, and frame-level embeddings.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="keyframes")# 1. A bucket for the videos, and a collection that cuts at scene changes and# writes a thumbnail and a multimodal embedding for every scenebucket = client.buckets.create(bucket_name="product-videos",bucket_schema={"properties": {"video": {"type": "video"}}},)collection = client.collections.create(collection_name="product-videos",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"split_method": "scene","scene_detection_threshold": 0.5,"enable_thumbnails": True,},},)# 2. Upload and processclient.buckets.upload(bucket["bucket_id"],blobs=[{"property": "video", "type": "video", "data": "s3://your-bucket/product-videos/unboxing.mp4"}],)client.collections.trigger(collection["collection_id"])# 3. Search the scenes visuallyretriever = client.retrievers.create(retriever_name="keyframe-search",collection_identifiers=["product-videos"],input_schema={"query": {"type": "text", "required": True}},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding","query": {"input_mode": "text", "value": "{{INPUT.query}}"},"top_k": 20,},],"final_top_k": 20,},},],)results = client.retrievers.execute(retriever["retriever_id"], inputs={"query": "product being unboxed"})for doc in results["documents"]:print(doc["thumbnail_url"], doc["start_time"], doc["score"])
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Processing a Large Media Library
Run extraction over tens of thousands of files without paying for the same object twice. Price the job before you submit it, keep new uploads flowing in on their own, and move cold collections out of the vector store while they stay searchable.
Text Grouping
Group video segments based on unique text appearing on screen
Object Grouping
Segment and group objects across video frames
Scene Splitting
Detect and segment distinct scenes in video content
Face Grouping
Detect, track, and group faces across video frames