NEWVectors or files. Pick a path.Start →
    Training

    Video Transcription & Indexing Pipeline

    Automatically transcribe video content with speaker identification, timestamps, and full-text indexing for downstream search and analytics.

    video
    audio
    text
    Single Tier
    from mixpeek import Mixpeek
    client = Mixpeek(api_key="YOUR_API_KEY", namespace="transcripts")
    # 1. A bucket for the recordings, and a collection that cuts each one at pauses
    # and transcribes and embeds every segment
    bucket = client.buckets.create(
    bucket_name="meeting-recordings",
    bucket_schema={"properties": {"recording": {"type": "video"}}},
    )
    collection = client.collections.create(
    collection_name="meetings",
    source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
    feature_extractor={
    "feature_extractor_name": "multimodal_extractor",
    "version": "v1",
    "parameters": {
    "split_method": "silence",
    "run_transcription": True,
    "run_transcription_embedding": True,
    },
    },
    )
    # 2. Upload and process. The managed extractor does not label speakers, so
    # diarization runs in your own pipeline when you need it.
    client.buckets.upload(
    bucket["bucket_id"],
    blobs=[{"property": "recording", "type": "video", "data": "s3://your-bucket/meeting-recordings/weekly-sync.mp4"}],
    )
    client.collections.trigger(collection["collection_id"])
    # 3. Read the transcript one segment at a time, with its time range
    docs = client.documents.list(collection["collection_id"], page_size=50)
    for doc in docs["results"]:
    print(doc["start_time"], doc["end_time"], doc["transcription"])

    Feature Extractors

    Multimodal Extractor

    Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.

    Retriever Stages

    Related Recipes & Resources

    Explore these related resources to deepen your understanding and discover more powerful features