Dense Search Over Your Own Embeddings, and What Hybrid Needs
Upsert documents you embedded elsewhere into an MVS namespace and search them by raw vector through the features search endpoint. Hybrid BM25 plus dense is not part of a plain BYO upsert: the documents carry dense vectors only and no text index is created. If you want a lexical leg later, declare a TEXT payload index on the field when you create the namespace; this recipe shows that declaration and the dense search that works today.
"FastAPI Pydantic v2 validation patterns"
Why This Matters
Pure vector search misses exact identifiers and error codes, and teams that bring their own embeddings often assume keyword matching comes with the vector store. On BYO documents it does not unless the text index exists before the first upsert. Declaring it up front is cheap; discovering its absence after indexing a corpus is not.
import requestsfrom openai import OpenAIfrom mixpeek import Mixpeekopenai = OpenAI(api_key="YOUR_OPENAI_KEY")client = Mixpeek(api_key="YOUR_API_KEY")NAMESPACE = "byo-hybrid"HEADERS = {"Authorization": "Bearer YOUR_API_KEY"}def embed(text):return openai.embeddings.create(model="text-embedding-3-small", input=text).data[0].embedding# 1. A standalone namespace for your vectorsclient.namespaces.create(namespace_id=NAMESPACE,mode="standalone",vector_configs=[{"name": "dense", "dimension": 1536, "metric": "cosine"}],)# 2. BM25 reads the namespace's text payload indexes, so declare one on the field that# holds the text before the first upsert. The SDK has no namespace update, so this is REST.requests.patch(f"https://api.mixpeek.com/v1/namespaces/{NAMESPACE}",headers=HEADERS,json={"payload_indexes": [{"field_name": "content", "type": "text"}]},)# 3. Upsert vectors with the text in the indexed payload fielddocuments = ["FastAPI uses Pydantic v2 for data validation and serialization","Express.js middleware handles request and response transformations","Django ORM provides database abstraction with the QuerySet API",]client.namespaces.documents.upsert(namespace_id=NAMESPACE,documents=[{"document_id": f"doc-{i}", "vectors": {"dense": embed(text)}, "payload": {"content": text}}for i, text in enumerate(documents)],)# 4. One stage runs both legs: dense on the vector, lexical on the text, fused with RRFretriever = client.retrievers.create(retriever_name="byo-hybrid-search",input_schema={"qv": {"type": "array", "required": True}, "query": {"type": "text", "required": True}},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "dense","query": {"input_mode": "vector", "value": "{{INPUT.qv}}"},"top_k": 20,},{"feature_uri": "dense","query": {"input_mode": "text", "value": "{{INPUT.query}}"},"top_k": 20,"lexical": True,},],"fusion": "rrf","final_top_k": 5,},},],)query = "FastAPI Pydantic validation"results = client.retrievers.execute(retriever["retriever_id"], inputs={"qv": embed(query), "query": query})for doc in results["documents"]:print(round(doc["score"], 3), doc.get("content"))
Feature Extractors
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
Documentation
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Document Intelligence Search
Extract and search through PDFs, presentations, and documents. Combines OCR, layout analysis, and semantic search for comprehensive document retrieval.
BYO Embeddings Vector Search
Bring pre-computed embeddings from any provider (OpenAI, Cohere, Together, etc.) and upsert them directly into MVS for instant vector search. No feature extractors, no pipelines -- just embeddings in, results out.
RAG with MVS Standalone
Complete RAG pipeline using MVS for retrieval and OpenAI for generation. Chunk your documents, embed them with any provider, store in MVS, retrieve relevant context, and generate answers -- no managed feature extractors needed.
Web Scraper
Extract structured data from webpages while maintaining semantic context and relationships
Text Embedding
Extract semantic embeddings from documents, transcripts and text content
Named Entity Recognition
Identify and extract named entities like people, organizations, and locations