Audio & Podcast Search Pipeline
Make audio content searchable by transcribing and embedding spoken content. Find specific moments in podcasts, calls, and recordings.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="audio-search")# 1. A bucket for the episodes, and a collection that cuts audio at pauses and transcribes and embeds each segmentbucket = client.buckets.create(bucket_name="podcasts",bucket_schema={"properties": {"episode": {"type": "audio",},},},)collection = client.collections.create(collection_name="podcasts",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"split_method": "silence","run_transcription": True,"run_transcription_embedding": True,},},)# 2. Upload and processclient.buckets.upload(bucket["bucket_id"],blobs=[{"property": "episode", "type": "audio", "data": "s3://your-bucket/podcasts/episode-212.mp3"}],)client.collections.trigger(collection["collection_id"])# 3. Semantic search over the transcript embeddingsretriever = client.retrievers.create(retriever_name="podcast-search",collection_identifiers=["podcasts"],input_schema={"query": {"type": "text","required": True,},},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/multilingual_e5_large_instruct_v1","query": {"input_mode": "text","value": "{{INPUT.query}}",},"top_k": 20,},],"final_top_k": 20,},},],)# 4. Searchresults = client.retrievers.execute(retriever["retriever_id"],inputs={"query": "discussion about AI regulation in Europe",},)for doc in results["documents"]:print(doc["start_time"], doc["transcription"], doc["score"])
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
Use Cases Using This Recipe
Earnings Call Signal Extraction
Extract predictive audio and text signals from earnings calls at scale
Text + audio + video (vs. text-only)
Feature modality coverage
Quantitative hedge funds, systematic trading desks, and fundamental research teams analyzing 500+ earnings events per quarter
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Audio Embedding
Extract semantic embeddings from audio content for similarity search
Speech to Text
Convert speech content to text with timestamps and confidence scores
Audio Classification
Classify audio content into categories like music, speech, noise, etc.
Speaker Diarization
Identify and separate different speakers in audio content
Audio Event Detection
Detect specific audio events like gunshots, glass breaking, alarms, etc.