Video Scene Search
Find specific scenes within videos using natural language descriptions. The pipeline detects scene boundaries, generates embeddings for each scene, and enables precise timestamp-level search across an entire video library. Query for visual content, actions, or spoken dialogue.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="video-scenes")# 1. A bucket for the videos, and a collection that detects scene boundaries and embeds each scenebucket = client.buckets.create(bucket_name="videos",bucket_schema={"properties": {"video": {"type": "video",},},},)collection = client.collections.create(collection_name="video_scenes",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"split_method": "scene","scene_detection_threshold": 0.5,"run_transcription": True,"run_transcription_embedding": True,},},)# 2. Upload and processclient.buckets.upload(bucket["bucket_id"],blobs=[{"property": "video", "type": "video", "data": "s3://your-bucket/videos/training-session.mp4"}],)client.collections.trigger(collection["collection_id"])# 3. Scene search over visual and transcript embeddings, then a rerank against each scene transcriptretriever = client.retrievers.create(retriever_name="scene-search",collection_identifiers=["video_scenes"],input_schema={"query": {"type": "text","required": True,},},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding","query": {"input_mode": "text","value": "{{INPUT.query}}",},"top_k": 50,},{"feature_uri": "mixpeek://multimodal_extractor@v1/multilingual_e5_large_instruct_v1","query": {"input_mode": "text","value": "{{INPUT.query}}",},"top_k": 50,},],"fusion": "rrf","final_top_k": 50,},},{"stage_name": "rerank","stage_id": "rerank","parameters": {"inference_name": "BAAI__bge_reranker_v2_m3","query": "{{INPUT.query}}","document_field": "transcription","top_k": 10,},},],)# 4. Searchresults = client.retrievers.execute(retriever["retriever_id"],inputs={"query": "person presenting a chart on a whiteboard",},)for doc in results["documents"]:print(doc["source_object_id"], doc["start_time"], doc["end_time"], doc["score"])
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
rerank
Rerank documents using cross-encoder models for accurate relevance
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Hierarchical Classification
Assign content to multi-level category hierarchies using embedding-based classification. Define your taxonomy once, then classify new content automatically with confidence scores.
Brand Safety & Ad Verification Pipeline
GARM-compliant brand safety pipeline for ad networks. Analyze video and image creatives for brand safety violations before serving.
Automated Video Tagging
Automatically generate descriptive tags for video content using scene analysis, object detection, and taxonomy classification. Each video receives structured labels for scenes, objects, actions, and custom business categories without manual annotation.
Searchable Video Library
Turn an unstructured video archive into a fully searchable library. Each video is decomposed into scenes with transcriptions, visual embeddings, and metadata. Users search by natural language and jump directly to the relevant moment in any video.
Semantic Multimodal Search
Unified semantic search across all content types. Query by natural language and retrieve relevant video clips, images, audio segments, and documents based on meaning-not keywords or manual tags.