Video Transcription & Indexing Pipeline
Automatically transcribe video content with speaker identification, timestamps, and full-text indexing for downstream search and analytics.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="transcripts")# 1. A bucket for the recordings, and a collection that cuts each one at pauses# and transcribes and embeds every segmentbucket = client.buckets.create(bucket_name="meeting-recordings",bucket_schema={"properties": {"recording": {"type": "video"}}},)collection = client.collections.create(collection_name="meetings",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"split_method": "silence","run_transcription": True,"run_transcription_embedding": True,},},)# 2. Upload and process. The managed extractor does not label speakers, so# diarization runs in your own pipeline when you need it.client.buckets.upload(bucket["bucket_id"],blobs=[{"property": "recording", "type": "video", "data": "s3://your-bucket/meeting-recordings/weekly-sync.mp4"}],)client.collections.trigger(collection["collection_id"])# 3. Read the transcript one segment at a time, with its time rangedocs = client.documents.list(collection["collection_id"], page_size=50)for doc in docs["results"]:print(doc["start_time"], doc["end_time"], doc["transcription"])
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Processing a Large Media Library
Run extraction over tens of thousands of files without paying for the same object twice. Price the job before you submit it, keep new uploads flowing in on their own, and move cold collections out of the vector store while they stay searchable.
Video Content Analytics Pipeline
Analyze video content at scale to extract insights: scene composition, speaker time, topic distribution, and sentiment across your video library.
Video RAG Pipeline
Retrieval-augmented generation specifically designed for video content. Decomposes videos into scenes and transcripts, retrieves relevant segments for a given question, and passes them as context to an LLM with precise timestamp citations.
Searchable Video Library
Turn an unstructured video archive into a fully searchable library. Each video is decomposed into scenes with transcriptions, visual embeddings, and metadata. Users search by natural language and jump directly to the relevant moment in any video.
Dataset Versioning
Treat versioned object storage as your dataset's source of truth. Capture complete snapshots-raw assets, embeddings, and cluster assignments-for deterministic reconstruction at any point in time.