NEWVectors or files. Pick a path.Start →
    embedding

    Mixed
    Embeddings
    Converter

    Generate unified vector embeddings from mixed-modality inputs -- text, images, audio, and video combined. Enables cross-modal search where any modality can query any other modality in a single vector space.

    Max file size: 5 GB
    Estimated: 1-15 sec depending on modality
    8 input formats

    How It Works

    1

    Provide one or more inputs of any modality.

    2

    Each input is processed through its modality-specific encoder.

    3

    Modality embeddings are projected into a shared vector space.

    4

    A fused embedding is produced that represents the combined input.

    5

    The unified embedding enables cross-modal similarity search.

    Code Examples

    import os, requests
    
    API = "https://api.mixpeek.com"
    H = {"Authorization": f"Bearer {os.environ['MIXPEEK_API_KEY']}",
         "X-Namespace": os.environ["NAMESPACE_ID"]}
    
    # 1. a bucket, with a schema that declares the field you will send
    bucket = requests.post(f"{API}/v1/buckets", headers=H, json={
        "bucket_name": "multimodal-inputs",
        "bucket_schema": {"properties": {"multimodal": {"type": "video"}}},
    }).json()
    
    # 2. land the file as an object. the URL goes in data, on the blob
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/objects", headers=H, json={
        "key_prefix": "run-1",
        "blobs": [{"property": "multimodal", "type": "video",
                   "data": "https://example.com/clip.mp4"}],
    })
    
    # 3. a collection over that bucket, running the extractor
    collection = requests.post(f"{API}/v1/collections", headers=H, json={
        "collection_name": "multimodal-to-embeddings",
        "source": {"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},
        "feature_extractor": {"feature_extractor_name": "universal_extractor", "version": "v1"},
    }).json()
    
    # 4. run extraction over the bucket
    requests.post(f"{API}/v1/buckets/{bucket['bucket_id']}/batches", headers=H, json={
        "collection_ids": [collection["collection_id"]],
        "auto_submit": True,
    })
    
    # 5. read the output
    docs = requests.get(
        f"{API}/v1/collections/{collection['collection_id']}/documents", headers=H
    ).json()
    print(docs)

    Use Cases

    Search videos using text queries and vice versa
    Build unified search across documents, images, and audio
    Create recommendation systems that span content types
    Enable 'find similar' features across an entire media library

    Supported Input Formats

    JPEG
    PNG
    MP4
    MP3
    WAV
    TXT
    PDF
    JSON

    Quick Info

    Categoryembedding
    Max File Size5 GB
    Est. Time1-15 sec depending on modality

    Processing millions of files?

    Run this as a managed pipeline over your whole library, no infrastructure to build or maintain. Talk to us about processing at scale.

    Run it over a library

    Mixpeek runs this conversion as a pipeline over a whole library in your object storage, with the output landing as queryable documents. It is not a single-file converter.

    Frequently Asked Questions

    Ready to convert mixed to embeddings?

    Start using the Mixpeek Multimodal to Embeddings in minutes. Sign up for a free API key and follow the documentation to get started.