Video RAG Pipeline
Retrieval-augmented generation specifically designed for video content. Decomposes videos into scenes and transcripts, retrieves relevant segments for a given question, and passes them as context to an LLM with precise timestamp citations.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="training-videos")# 1. Training videos are cut at pauses, transcribed, and each passage is embeddedbucket = client.buckets.create(bucket_name="training-videos",bucket_schema={"properties": {"video": {"type": "video"}}},)collection = client.collections.create(collection_name="training-videos",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"split_method": "silence","run_transcription": True,"run_transcription_embedding": True,},},)client.buckets.upload(bucket["bucket_id"],blobs=[{"property": "video", "type": "video", "data": "s3://your-bucket/training/firewall-setup.mp4"}],)client.collections.trigger(collection["collection_id"])# 2. The retriever finds passages, reranks them and writes the answerretriever = client.retrievers.create(retriever_name="video-rag",collection_identifiers=["training-videos"],input_schema={"query": {"type": "text", "required": True}},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/multilingual_e5_large_instruct_v1","query": {"input_mode": "text", "value": "{{INPUT.query}}"},"top_k": 50,},],"final_top_k": 50,},},{"stage_name": "rerank","stage_id": "rerank","parameters": {"inference_name": "BAAI__bge_reranker_v2_m3","query": "{{INPUT.query}}","document_field": "transcription","top_k": 8,},},{"stage_name": "answer","stage_id": "summarize","parameters": {"prompt": "Answer the question {{INPUT.query}} using only these numbered passages, and cite the passage numbers. {{DOCUMENTS}}","provider": "google","model_name": "gemini-2.5-flash-lite","content_field": "transcription","output_field": "answer","include_sources": True,},},],)results = client.retrievers.execute(retriever["retriever_id"], inputs={"query": "How do I configure the firewall settings?"})answer = results["documents"][0]print(answer["answer"])# The summary document lists the documents it read; fetch them for citationsfor i, document_id in enumerate(answer.get("source_document_ids") or [], 1):source = client.documents.get(collection["collection_id"], document_id)print(f"[{i}]", source.get("root_object_id"), source.get("start_time"), source.get("end_time"))
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
rerank
Rerank documents using cross-encoder models for accurate relevance
summarize
Condense multiple documents into a summary using an LLM
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Video Transcription & Indexing Pipeline
Automatically transcribe video content with speaker identification, timestamps, and full-text indexing for downstream search and analytics.
Video Content Analytics Pipeline
Analyze video content at scale to extract insights: scene composition, speaker time, topic distribution, and sentiment across your video library.
Searchable Video Library
Turn an unstructured video archive into a fully searchable library. Each video is decomposed into scenes with transcriptions, visual embeddings, and metadata. Users search by natural language and jump directly to the relevant moment in any video.
Semantic Multimodal Search
Unified semantic search across all content types. Query by natural language and retrieve relevant video clips, images, audio segments, and documents based on meaning-not keywords or manual tags.
Feature Extraction
Multi-tier feature extraction that decomposes content into searchable components: embeddings, transcripts, detected objects, OCR text, scene boundaries, and more. The foundation for all downstream retrieval and analysis.