Multimodal Search with MVS
Build multimodal search by embedding different content types (text, images, video frames) with your own models and searching across them in a single MVS namespace. Use CLIP or any multimodal embedding model for cross-modal retrieval.
"dog playing outdoors"
Why This Matters
Multimodal search without a managed pipeline. Use your preferred CLIP, SigLIP, or any embedding model to embed text, images, and video into a shared vector space, then search across all of them with MVS.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY")NAMESPACE = "multimodal-search"# --- local embedding model (runs on your machine; any CLIP-compatible model works) ---import open_clipimport torchfrom PIL import Imagemodel, _, preprocess = open_clip.create_model_and_transforms("ViT-B-32", pretrained="laion2b_s34b_b79k")tokenizer = open_clip.get_tokenizer("ViT-B-32")def embed_text(text):with torch.no_grad():features = model.encode_text(tokenizer([text]))return (features / features.norm(dim=-1, keepdim=True))[0].tolist()def embed_image(path):with torch.no_grad():features = model.encode_image(preprocess(Image.open(path)).unsqueeze(0))return (features / features.norm(dim=-1, keepdim=True))[0].tolist()# --- end local embedding model ---# 1. A standalone namespace for 512-number CLIP vectorsclient.namespaces.create(namespace_id=NAMESPACE,mode="standalone",vector_configs=[{"name": "clip", "dimension": 512, "metric": "cosine"}],)# 2. Text and images share the space, so both go in as the same vector nametexts = ["A golden retriever playing fetch in the park", "Sunset over the ocean with sailboats"]images = ["dog_park.jpg", "sunset.jpg"]client.namespaces.documents.upsert(namespace_id=NAMESPACE,documents=[{"document_id": f"text-{i}", "vectors": {"clip": embed_text(t)}, "payload": {"modality": "text", "content": t}}for i, t in enumerate(texts)]+ [{"document_id": f"image-{i}", "vectors": {"clip": embed_image(p)}, "payload": {"modality": "image", "file": p}}for i, p in enumerate(images)],)# 3. A retriever over the CLIP vectorretriever = client.retrievers.create(retriever_name="clip-search",input_schema={"qv": {"type": "array", "required": True}},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "clip","query": {"input_mode": "vector", "value": "{{INPUT.qv}}"},"top_k": 10,},],"final_top_k": 10,},},],)# A text query finds images and text alikequery = embed_text("dog playing outdoors")results = client.retrievers.execute(retriever["retriever_id"], inputs={"qv": query})for doc in results["documents"]:print(round(doc["score"], 3), doc.get("modality"), doc.get("content") or doc.get("file"))# The same query, filtered to imagesimages_only = client.retrievers.execute(retriever["retriever_id"],inputs={"qv": query},filters={"AND": [{"field": "modality", "operator": "eq", "value": "image"}]},)print(len(images_only["documents"]), "image results")
Feature Extractors
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
Documentation
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Hybrid Search Pipeline
Combine vector search with keyword search (BM25) across text, images, and video for the most comprehensive multimodal retrieval system.
Document Intelligence Search
Extract and search through PDFs, presentations, and documents. Combines OCR, layout analysis, and semantic search for comprehensive document retrieval.
BYO Embeddings Vector Search
Bring pre-computed embeddings from any provider (OpenAI, Cohere, Together, etc.) and upsert them directly into MVS for instant vector search. No feature extractors, no pipelines -- just embeddings in, results out.
RAG with MVS Standalone
Complete RAG pipeline using MVS for retrieval and OpenAI for generation. Chunk your documents, embed them with any provider, store in MVS, retrieve relevant context, and generate answers -- no managed feature extractors needed.
Multimodal RAG Pipeline
Build a retrieval-augmented generation system that works with text, images, and video. Feed relevant multimodal context to LLMs for grounded responses.
Taxonomy Enrichment Pipeline
Automatically classify and tag content using custom taxonomies. Map your content to IAB categories, custom hierarchies, or industry-specific classifications.