NEWVectors or files. Pick a path.Start →
    media

    Video
    Embeddings
    Converter

    Generate dense vector embeddings for video content using multimodal models. Embeddings capture visual, audio, and temporal features, enabling semantic search and similarity matching across video collections.

    Max file size: 5 GB
    Estimated: 3-12 min per hour of video
    5 input formats

    How It Works

    1

    Upload your video or provide a URL.

    2

    The video is segmented into clips based on scene boundaries.

    3

    Each clip is processed through a multimodal embedding model (CLIP, SigLIP, or E5).

    4

    Audio and visual features are fused into a single embedding per segment.

    5

    Embeddings are returned as float arrays ready for vector indexing.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "video-inputs",
        "bucket_schema": {"properties": {"video": {"type": "video"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "video", "type": "video",
                   "data": "https://example.com/clip.mp4"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "video-to-embeddings",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "multimodal_extractor", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Build semantic video search engines
    Detect near-duplicate or pirated video content
    Cluster similar videos for recommendation systems
    Enable cross-modal retrieval (search videos with text queries)

    Supported Input Formats

    MP4
    MOV
    AVI
    MKV
    WebM

    Quick Info

    Categorymedia
    Max File Size5 GB
    Est. Time3-12 min per hour of video

    Processing millions of hours of video?

    Run this as a managed pipeline over your entire video archive, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.

    Frequently Asked Questions

    Ready to convert video to embeddings?

    Start using the Mixpeek Video to Embeddings in minutes. Sign up for a free API key and follow the documentation to get started.