NEWMVS for embeddings. Managed for files. Both on object storage.Vectors or files. Pick a path.Start →

Start here

Vector Store (MVS)

Bring your own vectors. Dense, sparse, and BM25 search on object storage.

Managed Indexing

Connect a bucket and auto-extract scenes, faces, OCR, transcripts, and embeddings.

Build

Compose multi-stage search in <100ms: filter, join, rerank.

Feature Extractors

Typed pipelines for faces, scenes, transcripts, OCR, fingerprints.

S3, GCS, R2, Mux, LangChain, MCP, and more. Connect your stack.

Generate and store embeddings from 50+ models, then search them.

By Industry

Map, search, and reuse the moments that perform. Plugs into iconik & Mux.

Talent search, brand safety, creative analytics.

Scene search, recommendation, archive access.

Visual search, PDP enrichment, catalog QA.

Lecture search, transcript Q&A, content safety.

View all solutions →

By Use Case

Face & Person Search

Find anyone across video libraries in milliseconds.

IP & Copyright Detection

Logos, songs, faces: one pipeline, one report.

Visual Taste & Recs

Scene-similarity ranked recommendations with RL.

Brand & Ad Safety

Pre-publish content screening at bid-time speeds.

View all use cases →

Build

API reference, SDKs, recipes, and architecture guides.

Launches, deep dives, and field notes from our engineers.

Browse supported HuggingFace models by task and modality.

See what teams are building with Mixpeek.

Education

Vendor-neutral deep dives on perception, retrieval, and embeddings.

Multimodal University

Fundamentals of multimodal retrieval, modules + certs.

Every term you need: embeddings to re-rankers.

Talks, demos, and customer sessions on demand.

Mixpeek vs. Pinecone, Weaviate, Twelve Labs, more.

Mission, team, and the multimodal vision.

We're hiring across research, infra, and design.

Talk to sales, support, or press.

White-glove 30-day production pilot for new customers.

Sign in Request Demo Get started →

Back to Glossary

What is Semantic Join

Semantic Join - A cross-collection enrichment operation that attaches context from one collection to results from another, using semantic similarity as the join key.

A semantic join is the multimodal equivalent of a SQL JOIN. In structured databases, JOINs combine rows from different tables using foreign keys. In a multimodal data warehouse, enrich stages combine results from different collections using embedding similarity or document relationships. This enables cross-referencing without pre-defined foreign keys.

How It Works

After a retrieval pipeline produces results from one collection (e.g., media library search), an enrich stage queries a second collection (e.g., brand safety scores) to attach contextual data to each result. The join can be by document ID, semantic similarity, or metadata matching.

Examples

Search media library for celebrity appearances → enrich with brand safety scores from a separate collection
Find similar products → enrich with pricing and availability from a catalog collection
Detect copyrighted audio → enrich with licensing terms from a rights database
Find relevant document passages → enrich with author and classification metadata

Best Practices

Use enrich stages after reduce stages to minimize the number of cross-collection lookups
Keep enrichment collections focused: one collection per enrichment type (brand scores, rights, metadata)
Use semantic joins for fuzzy matching and document_enrich for exact ID-based joins

Related Pages

Document Enrich stage: /docs/retrieval/stages/document-enrich
Retrieval Cookbook: /docs/retrieval/cookbook
Blog: Multi-Stage Retrieval Pipelines - /blog/multi-stage-retrieval-pipelines

Put it to work: search your own files, free

Managed Mixpeek

Put multimodal search to work

Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

Start with Managed

MVS · bring your own

Already have vectors?

Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

Building an agent? Connect Mixpeek over MCP

Related Terms

ACID API Blob Storage CLIP Embedding