Multimodal RAG
Retrieval-augmented generation across video, images, and text. Retrieve relevant multimodal context, then pass to your LLM with citations back to source timestamps and frames.
"How did the product launch go? Cite specific video clips and document timestamps"
Why This Matters
RAG quality depends on retrieval quality. Mixpeek handles the multimodal retrieval infrastructure while you bring your preferred generation model.
from openai import OpenAIfrom mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="launch-recordings")openai = OpenAI(api_key="YOUR_OPENAI_KEY")# 1. Recordings are cut at pauses, transcribed, and each passage is embeddedbucket = client.buckets.create(bucket_name="launch-recordings",bucket_schema={"properties": {"recording": {"type": "video"}}},)collection = client.collections.create(collection_name="launch-recordings",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"split_method": "silence","run_transcription": True,"run_transcription_embedding": True,},},)client.buckets.upload(bucket["bucket_id"],blobs=[{"property": "recording", "type": "video", "data": "s3://your-bucket/launch/customer-call.mp4"}],)client.collections.trigger(collection["collection_id"])# 2. Retrieve passages, then rerank them against the questionretriever = client.retrievers.create(retriever_name="launch-context",collection_identifiers=["launch-recordings"],input_schema={"query": {"type": "text", "required": True}},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/multilingual_e5_large_instruct_v1","query": {"input_mode": "text", "value": "{{INPUT.query}}"},"top_k": 50,},],"final_top_k": 50,},},{"stage_name": "rerank","stage_id": "rerank","parameters": {"inference_name": "BAAI__bge_reranker_v2_m3","query": "{{INPUT.query}}","document_field": "transcription","top_k": 8,},},],)question = "How did the product launch go?"results = client.retrievers.execute(retriever["retriever_id"], inputs={"query": question})# 3. Number each passage with its source and timestamp so the answer can cite itcontext = " ".join(f"[{i}] {doc['transcription']} (source {doc['root_object_id']} at {doc['start_time']}s)."for i, doc in enumerate(results["documents"], 1))response = openai.chat.completions.create(model="gpt-4o",messages=[{"role": "system", "content": "Answer from the numbered context and cite the numbers. Context: " + context},{"role": "user", "content": question},],)print(response.choices[0].message.content)
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
rerank
Rerank documents using cross-encoder models for accurate relevance
Documentation
Use Cases Using This Recipe
Course Content Intelligence
Make every lecture moment searchable and actionable
80% reduction
Content discovery time
EdTech platforms, universities, and corporate L&D teams managing 1,000+ hours of educational content
Epstein Files Intelligence
Search and analyze thousands of declassified legal documents
100% of corpus indexed
Document searchability
Investigative journalists, legal researchers, OSINT analysts, and public interest organizations working with large declassified document sets
Government Intelligence
Multimodal search and analysis for government document repositories
100% unified index
Cross-department search coverage
Government agencies, policy researchers, compliance teams, and public affairs professionals managing multi-department document repositories
Semantic Search for Knowledge Bases
Find answers by meaning, not keywords, across your entire knowledge repository
85% of queries answered on first search vs. 40% baseline
First-search success rate
Knowledge management teams, internal documentation owners, customer support organizations, and EdTech platforms maintaining 10K+ articles, documents, and multimedia resources
Enterprise RAG Search
Ask questions across all your enterprise data and get sourced, verifiable answers
80% faster from question to answer
Information retrieval time
Financial services firms, consulting organizations, legal teams, and enterprise knowledge workers who need to synthesize information across thousands of internal documents, reports, and presentations
Multimodal RAG
Retrieval-augmented generation that understands text, images, video, and audio
+26% over text-only RAG
Answer accuracy with multimodal context
AI engineering teams, product builders, and enterprise developers building RAG applications that need to reason over documents, images, video, and audio rather than text alone
AI Compliance Document Review
Automate regulatory document review with multimodal AI understanding
10x faster
Review cycle time
Compliance teams, regulatory affairs departments, and legal operations groups reviewing 1,000+ regulatory documents per quarter across banking, insurance, pharma, and financial services
AI-Powered E-Discovery
Accelerate legal discovery across documents, emails, video, and audio evidence
83% lower cost per document
Review cost reduction
Law firms, corporate legal departments, litigation support providers, and e-discovery vendors processing 100,000+ documents per matter across civil litigation, regulatory investigations, and internal compliance reviews
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Semantic Join
Bridge extracted content features with business reference data. Join video clips to product catalogs, detected faces to employee directories, or documents to compliance frameworks-all via embedding similarity.
Semantic Multimodal Search
Unified semantic search across all content types. Query by natural language and retrieve relevant video clips, images, audio segments, and documents based on meaning-not keywords or manual tags.
Feature Extraction
Multi-tier feature extraction that decomposes content into searchable components: embeddings, transcripts, detected objects, OCR text, scene boundaries, and more. The foundation for all downstream retrieval and analysis.
Clustering & Theme Discovery
Unsupervised clustering that groups content into semantic themes using HDBSCAN. Surfaces hidden patterns, content variants, and outliers without requiring predefined labels.
Anomaly Detection
Identify outliers and anomalous content using embedding distance from cluster centroids. Flag quality issues, novel content, or items that don't match expected patterns.