How do I convert text, images, video and audio into embeddings?
Run the content through an embedding model and keep the list of numbers it returns. Text goes to a text model and images to a vision-language model, while video and audio are cut into segments that are embedded one at a time. Mixpeek runs that step over the files in your object storage, stores the vectors in an index you can search, and records which model wrote each one.
How do I turn one piece of text into an embedding?
Load an embedding model and call it on the text. The open-source sentence-transformers library needs a few lines, and the result is one vector per string: 1,024 numbers for multilingual E5 Large Instruct, the model Mixpeek's text extractor runs. Store each vector next to an ID for its text, and a search compares the query's vector with the stored ones.
Embed the query with the same model you used for the stored text. A vector from one model cannot be compared with a vector from another, even when both have the same length. E5 instruct models also expect a one-line task instruction in front of a query, while passages go in as plain text.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/multilingual-e5-large-instruct")
passages = [
"Replacement filter cartridge, part AB-4471-X",
"Trail running shoe with a waterproof upper",
]
vectors = model.encode(passages, normalize_embeddings=True)
print(vectors.shape) # (2, 1024)Comparing open models before you pick one? See the self-hosted embedding models comparison, the multimodal embedding models comparison and the model hub.
What is an embedding, and what makes it multimodal?
An embedding is a list of numbers a model produces for a piece of content, arranged so that content with similar meaning gets similar numbers. A multimodal model does this for more than one kind of content in the same space, so a text query can be compared with an image or a video segment.
Text Embeddings
A text model turns a sentence or a passage into one vector. Mixpeek's text extractor runs multilingual E5 Large Instruct, which returns 1,024 numbers per chunk and covers more than 100 languages.
Image Embeddings
A vision-language model puts an image and a text query in the same space, so words can find pictures. The image extractor runs SigLIP over images and PDF pages and returns 768 numbers for each.
Video Embeddings
Video is cut into segments before anything is embedded. Each segment gets its own vector, with its transcript and on-screen text kept beside it, so a search lands on the moment inside the file.
Which embedding model should I use for each kind of content?
These are the models Mixpeek's extractors run, with the vector size each one returns, checked against the live extractor catalog in September 2026.
| Model | Modalities | Dimensions | Best For | Extractor |
|---|---|---|---|---|
| multilingual-e5-large-instruct | Text, in more than 100 languages | 1024 | Semantic text search and RAG over documents in many languages | Text |
| SigLIP base patch16-224 | Images, PDF pages, text queries | 768 | Finding images by a description or by an example image | Image |
| Vertex multimodal embedding | Video segments, images, text | 1408 | Searching video segments and images with a text query | Multimodal v1 |
| Gemini Embedding 2 | Video, audio, images, text | 3072 | One space for every file type in a library | Multimodal v2 |
| CLAP HTSAT-tiny | Audio, or the audio track of a video | 512 | Matching a sound, jingle or music clip to other audio | Audio fingerprint |
| ArcFace (SCRFD detection) | Faces in images, video and PDFs | 512 | Finding the same person across photos and footage | Face identity |
Already have vectors from your own model? Store and search them directly in Mixpeek Vector Store, and see pricing for what extraction and storage cost.
How do I generate embeddings for a whole library?
From files in object storage to searchable vectors in four steps.
Choose Models
Pick an extractor for each kind of file: text, image, multimodal for video and audio, or faces. Enterprise deployments can also run a model you bring as a custom extractor.
Ingest Content
Connect a bucket to your object storage or upload through the API. Each collection runs its extractor over the files in that bucket.
Generate Vectors
Extraction runs as batch jobs that split the files, call the model and retry failures, so a whole library goes through the same path as a single file.
Index & Search
Vectors are stored in Mixpeek Vector Store on object storage. A retriever searches one or more of those indexes and can add filters, keyword matching and reranking.
What happens to my embeddings when I change models?
Every vector belongs to the model that created it, so a new model means a new index. Mixpeek keeps each extractor's vectors in their own collection and leaves the source files in your bucket, which turns a model change into a backfill.
The Lock-In Problem
CLIP vectors and SigLIP vectors are not interchangeable. They live in completely different mathematical spaces. Mixing them in the same index silently degrades retrieval quality.
The Upgrade Trap
When a better model ships, the choice is to stay on the old one or re-encode everything. Re-encoding costs compute in proportion to the library, which is why most teams postpone it until it becomes urgent.
How Mixpeek Solves It
Every index records the extractor and version that wrote it. Run the old and new collections side by side, compare them, and switch the retriever when the new one is ready. The source files stay in your bucket for the next re-encode.
Upgrade without downtime
One collection per extractor version
Each extractor version writes its own collection. The old index keeps serving while the new one backfills, and the two sets of vectors are never mixed in one index.
Automatic re-encoding
Create a new collection with the updated extractor, point it at the same bucket, and trigger reprocessing. The batch pipeline handles backfill automatically.
Retriever-level cutover
Change the collection a retriever searches and the application keeps calling the same retriever. The switch is a configuration change.
Source-data retention
The source files stay in the bucket they came from, so the index can be rebuilt with a new model whenever you choose.
What can I build once my content is embedded?
The same vectors serve search, deduplication, classification, recommendations and retrieval for LLM answers.
Semantic Search
Search by meaning rather than keywords. Find relevant content even when the query uses different terminology than the source material.
Cross-Modal Retrieval
Query in one modality and retrieve results in another. Search video with text, find images with audio descriptions, or match documents to visual content.
Duplicate Detection
Identify near-duplicate content across your corpus by comparing embedding similarity. Works across modalities and catches semantic duplicates as well as pixel-identical copies.
Content Classification
Classify content into categories using embedding similarity to reference examples. Enable zero-shot classification without collecting labeled training data for each category.
Recommendation Systems
Build content recommendations by finding embeddings similar to user interaction history. Works across content types for multimodal recommendation.
RAG Applications
Power retrieval-augmented generation by embedding your knowledge base and retrieving relevant context for LLM prompts across text, images, and documents.
How do I generate and search embeddings with Mixpeek?
One call returns a single vector. A retriever searches everything a collection has already embedded.
import os
import requests
from mixpeek import Mixpeek
API_KEY = os.environ["MIXPEEK_API_KEY"]
# One vector from the model the text extractor indexes with
resp = requests.post(
"https://api.mixpeek.com/v1/inference",
headers={"Authorization": f"Bearer {API_KEY}", "X-Namespace": "my-namespace"},
json={
"feature_uri": "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1",
"inputs": {"text": "quarterly revenue growth in emerging markets"},
},
)
vector = resp.json()["data"]["embeddings"][0]
print(len(vector)) # 1024
# Search everything a collection has already embedded
client = Mixpeek(api_key=API_KEY, namespace="my-namespace")
results = client.retrievers.execute(
"ret_your_retriever_id",
inputs={"query": "product packaging with sustainability labels"},
)
for doc in results["documents"]:
print(doc["document_id"], doc["score"])Frequently Asked Questions
How do I convert text into an embedding in Python?
Load an embedding model and call it on the text. With the open-source sentence-transformers library, SentenceTransformer("intfloat/multilingual-e5-large-instruct").encode(["your text"]) returns one vector of 1,024 numbers per string. Mixpeek's inference endpoint returns the same model's vector over HTTP, and a collection produces one for every file in a bucket.
What are multimodal embeddings?
Multimodal embeddings are vectors that represent the meaning of more than one kind of content, such as text, images, video and audio, in a single shared space. Because matching content lands close together in that space, a text query can be compared directly with an image or a video segment.
What embedding models does Mixpeek run?
The extractors run multilingual E5 Large Instruct for text (1,024 dimensions), SigLIP for images and PDF pages (768), Vertex multimodal embedding (1,408) and Gemini Embedding 2 (3,072) for video, audio, images and text, CLAP for audio (512) and ArcFace for faces (512). Enterprise deployments can register another model as a custom extractor.
Can I use my own embedding models with Mixpeek?
Yes. On Enterprise deployments a custom extractor wraps your model and declares its input and output schema, including the vector size, and Mixpeek runs it with the same batching, retries and indexing as the built-in extractors. If you already have vectors, a standalone namespace in Mixpeek Vector Store accepts them as they are.
How do cross-modal embeddings work?
Cross-modal embeddings come from models trained with a contrastive objective that pulls matching pairs together, for example a photo and its caption. CLIP and SigLIP learn this for images and text, and CLAP learns it for audio and text. After training, a text query vector can be compared directly against image or audio vectors to find matching content.
What vector dimensions does Mixpeek support?
The built-in extractors return 512, 768, 1,024, 1,408 or 3,072 dimensions depending on the model. A standalone namespace stores vectors of whatever size your own model produces, and Mixpeek Vector Store holds dense and sparse vectors.
How does Mixpeek generate embeddings for a large library?
Batch processing runs on Ray. When a collection is triggered, the engine splits the files into batches, runs the extractor over them in parallel and retries items that fail, and you can follow the job through the tasks API.
Can I store multiple embedding types per document?
Yes. Run more than one collection over the same bucket, one extractor each, and a retriever can search several of those indexes in one query and fuse the results. A product can have a text embedding of its description and a SigLIP embedding of its photo.
How do I choose the right embedding model for my use case?
Start from what people will type and what they expect back. Text questions over written content suit the text extractor. Finding pictures with words, or with another picture, is what SigLIP was trained for. When a query should land on a moment inside a video or a recording, use the multimodal extractor, and use the face extractor to find a person. If two choices look close, build a collection with each over the same files and compare them with a retriever evaluation.
Can I upgrade embedding models without re-indexing everything at once?
Yes. Create a new collection with the new extractor over the same bucket and let it backfill while the old collection keeps serving. When the new index is ready, change the collection your retriever searches. Vectors from the two models are never mixed in one index.
Are embeddings from different models compatible with each other?
No. Vectors from different models live in unrelated coordinate spaces even when their length matches. Multilingual E5 Large Instruct and BGE-M3 both return 1,024 dimensions, and comparing a vector from one with a vector from the other returns noise. Mixpeek keeps each extractor's vectors in their own index and records which extractor and version wrote them.