NEWVectors or files. Pick a path.Start →
    Updated September 2026

    How do I convert text, images, video and audio into embeddings?

    Run the content through an embedding model and keep the list of numbers it returns. Text goes to a text model and images to a vision-language model, while video and audio are cut into segments that are embedded one at a time. Mixpeek runs that step over the files in your object storage, stores the vectors in an index you can search, and records which model wrote each one.

    How do I turn one piece of text into an embedding?

    Load an embedding model and call it on the text. The open-source sentence-transformers library needs a few lines, and the result is one vector per string: 1,024 numbers for multilingual E5 Large Instruct, the model Mixpeek's text extractor runs. Store each vector next to an ID for its text, and a search compares the query's vector with the stored ones.

    Embed the query with the same model you used for the stored text. A vector from one model cannot be compared with a vector from another, even when both have the same length. E5 instruct models also expect a one-line task instruction in front of a query, while passages go in as plain text.

    text_to_embedding.py
    from sentence_transformers import SentenceTransformer
    
    model = SentenceTransformer("intfloat/multilingual-e5-large-instruct")
    
    passages = [
        "Replacement filter cartridge, part AB-4471-X",
        "Trail running shoe with a waterproof upper",
    ]
    vectors = model.encode(passages, normalize_embeddings=True)
    print(vectors.shape)  # (2, 1024)

    Comparing open models before you pick one? See the self-hosted embedding models comparison, the multimodal embedding models comparison and the model hub.

    What is an embedding, and what makes it multimodal?

    An embedding is a list of numbers a model produces for a piece of content, arranged so that content with similar meaning gets similar numbers. A multimodal model does this for more than one kind of content in the same space, so a text query can be compared with an image or a video segment.

    Text Embeddings

    A text model turns a sentence or a passage into one vector. Mixpeek's text extractor runs multilingual E5 Large Instruct, which returns 1,024 numbers per chunk and covers more than 100 languages.

    multilingual-e5-large-instruct
    Gemini Embedding 2

    Image Embeddings

    A vision-language model puts an image and a text query in the same space, so words can find pictures. The image extractor runs SigLIP over images and PDF pages and returns 768 numbers for each.

    SigLIP base patch16-224
    Vertex multimodal embedding

    Video Embeddings

    Video is cut into segments before anything is embedded. Each segment gets its own vector, with its transcript and on-screen text kept beside it, so a search lands on the moment inside the file.

    Vertex multimodal embedding
    Gemini Embedding 2

    Audio Embeddings

    Audio is split into short windows and each window is embedded. CLAP places sounds and their text descriptions in one 512-dimension space, and speech can also be searched through its transcript.

    CLAP HTSAT-tiny
    Gemini Embedding 2

    Which embedding model should I use for each kind of content?

    These are the models Mixpeek's extractors run, with the vector size each one returns, checked against the live extractor catalog in September 2026.

    ModelModalitiesDimensionsBest ForExtractor
    multilingual-e5-large-instructText, in more than 100 languages1024Semantic text search and RAG over documents in many languagesText
    SigLIP base patch16-224Images, PDF pages, text queries768Finding images by a description or by an example imageImage
    Vertex multimodal embeddingVideo segments, images, text1408Searching video segments and images with a text queryMultimodal v1
    Gemini Embedding 2Video, audio, images, text3072One space for every file type in a libraryMultimodal v2
    CLAP HTSAT-tinyAudio, or the audio track of a video512Matching a sound, jingle or music clip to other audioAudio fingerprint
    ArcFace (SCRFD detection)Faces in images, video and PDFs512Finding the same person across photos and footageFace identity

    Already have vectors from your own model? Store and search them directly in Mixpeek Vector Store, and see pricing for what extraction and storage cost.

    How do I generate embeddings for a whole library?

    From files in object storage to searchable vectors in four steps.

    1

    Choose Models

    Pick an extractor for each kind of file: text, image, multimodal for video and audio, or faces. Enterprise deployments can also run a model you bring as a custom extractor.

    2

    Ingest Content

    Connect a bucket to your object storage or upload through the API. Each collection runs its extractor over the files in that bucket.

    3

    Generate Vectors

    Extraction runs as batch jobs that split the files, call the model and retry failures, so a whole library goes through the same path as a single file.

    4

    Index & Search

    Vectors are stored in Mixpeek Vector Store on object storage. A retriever searches one or more of those indexes and can add filters, keyword matching and reranking.

    What happens to my embeddings when I change models?

    Every vector belongs to the model that created it, so a new model means a new index. Mixpeek keeps each extractor's vectors in their own collection and leaves the source files in your bucket, which turns a model change into a backfill.

    The Lock-In Problem

    CLIP vectors and SigLIP vectors are not interchangeable. They live in completely different mathematical spaces. Mixing them in the same index silently degrades retrieval quality.

    The Upgrade Trap

    When a better model ships, the choice is to stay on the old one or re-encode everything. Re-encoding costs compute in proportion to the library, which is why most teams postpone it until it becomes urgent.

    How Mixpeek Solves It

    Every index records the extractor and version that wrote it. Run the old and new collections side by side, compare them, and switch the retriever when the new one is ready. The source files stay in your bucket for the next re-encode.

    Upgrade without downtime

    One collection per extractor version

    Each extractor version writes its own collection. The old index keeps serving while the new one backfills, and the two sets of vectors are never mixed in one index.

    Automatic re-encoding

    Create a new collection with the updated extractor, point it at the same bucket, and trigger reprocessing. The batch pipeline handles backfill automatically.

    Retriever-level cutover

    Change the collection a retriever searches and the application keeps calling the same retriever. The switch is a configuration change.

    Source-data retention

    The source files stay in the bucket they came from, so the index can be rebuilt with a new model whenever you choose.

    What can I build once my content is embedded?

    The same vectors serve search, deduplication, classification, recommendations and retrieval for LLM answers.

    Semantic Search

    Search by meaning rather than keywords. Find relevant content even when the query uses different terminology than the source material.

    Cross-Modal Retrieval

    Query in one modality and retrieve results in another. Search video with text, find images with audio descriptions, or match documents to visual content.

    Duplicate Detection

    Identify near-duplicate content across your corpus by comparing embedding similarity. Works across modalities and catches semantic duplicates as well as pixel-identical copies.

    Content Classification

    Classify content into categories using embedding similarity to reference examples. Enable zero-shot classification without collecting labeled training data for each category.

    Recommendation Systems

    Build content recommendations by finding embeddings similar to user interaction history. Works across content types for multimodal recommendation.

    RAG Applications

    Power retrieval-augmented generation by embedding your knowledge base and retrieving relevant context for LLM prompts across text, images, and documents.

    How do I generate and search embeddings with Mixpeek?

    One call returns a single vector. A retriever searches everything a collection has already embedded.

    embeddings_example.py
    import os
    import requests
    from mixpeek import Mixpeek
    
    API_KEY = os.environ["MIXPEEK_API_KEY"]
    
    # One vector from the model the text extractor indexes with
    resp = requests.post(
        "https://api.mixpeek.com/v1/inference",
        headers={"Authorization": f"Bearer {API_KEY}", "X-Namespace": "my-namespace"},
        json={
            "feature_uri": "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1",
            "inputs": {"text": "quarterly revenue growth in emerging markets"},
        },
    )
    vector = resp.json()["data"]["embeddings"][0]
    print(len(vector))  # 1024
    
    # Search everything a collection has already embedded
    client = Mixpeek(api_key=API_KEY, namespace="my-namespace")
    results = client.retrievers.execute(
        "ret_your_retriever_id",
        inputs={"query": "product packaging with sustainability labels"},
    )
    for doc in results["documents"]:
        print(doc["document_id"], doc["score"])

    Frequently Asked Questions

    How do I convert text into an embedding in Python?

    Load an embedding model and call it on the text. With the open-source sentence-transformers library, SentenceTransformer("intfloat/multilingual-e5-large-instruct").encode(["your text"]) returns one vector of 1,024 numbers per string. Mixpeek's inference endpoint returns the same model's vector over HTTP, and a collection produces one for every file in a bucket.

    What are multimodal embeddings?

    Multimodal embeddings are vectors that represent the meaning of more than one kind of content, such as text, images, video and audio, in a single shared space. Because matching content lands close together in that space, a text query can be compared directly with an image or a video segment.

    What embedding models does Mixpeek run?

    The extractors run multilingual E5 Large Instruct for text (1,024 dimensions), SigLIP for images and PDF pages (768), Vertex multimodal embedding (1,408) and Gemini Embedding 2 (3,072) for video, audio, images and text, CLAP for audio (512) and ArcFace for faces (512). Enterprise deployments can register another model as a custom extractor.

    Can I use my own embedding models with Mixpeek?

    Yes. On Enterprise deployments a custom extractor wraps your model and declares its input and output schema, including the vector size, and Mixpeek runs it with the same batching, retries and indexing as the built-in extractors. If you already have vectors, a standalone namespace in Mixpeek Vector Store accepts them as they are.

    How do cross-modal embeddings work?

    Cross-modal embeddings come from models trained with a contrastive objective that pulls matching pairs together, for example a photo and its caption. CLIP and SigLIP learn this for images and text, and CLAP learns it for audio and text. After training, a text query vector can be compared directly against image or audio vectors to find matching content.

    What vector dimensions does Mixpeek support?

    The built-in extractors return 512, 768, 1,024, 1,408 or 3,072 dimensions depending on the model. A standalone namespace stores vectors of whatever size your own model produces, and Mixpeek Vector Store holds dense and sparse vectors.

    How does Mixpeek generate embeddings for a large library?

    Batch processing runs on Ray. When a collection is triggered, the engine splits the files into batches, runs the extractor over them in parallel and retries items that fail, and you can follow the job through the tasks API.

    Can I store multiple embedding types per document?

    Yes. Run more than one collection over the same bucket, one extractor each, and a retriever can search several of those indexes in one query and fuse the results. A product can have a text embedding of its description and a SigLIP embedding of its photo.

    How do I choose the right embedding model for my use case?

    Start from what people will type and what they expect back. Text questions over written content suit the text extractor. Finding pictures with words, or with another picture, is what SigLIP was trained for. When a query should land on a moment inside a video or a recording, use the multimodal extractor, and use the face extractor to find a person. If two choices look close, build a collection with each over the same files and compare them with a retriever evaluation.

    Can I upgrade embedding models without re-indexing everything at once?

    Yes. Create a new collection with the new extractor over the same bucket and let it backfill while the old collection keeps serving. When the new index is ready, change the collection your retriever searches. Vectors from the two models are never mixed in one index.

    Are embeddings from different models compatible with each other?

    No. Vectors from different models live in unrelated coordinate spaces even when their length matches. Multilingual E5 Large Instruct and BGE-M3 both return 1,024 dimensions, and comparing a vector from one with a vector from the other returns noise. Mixpeek keeps each extractor's vectors in their own index and records which extractor and version wrote them.

    Start Generating Multimodal Embeddings

    One API for text, image, video, and audio embeddings. Explore a public demo for free, choose a plan from $25/month for your own data, or talk to us about enterprise deployment.