NEWVectors or files. Pick a path.Start →
    Similar

    Dense Search Over Your Own Embeddings, and What Hybrid Needs

    Upsert documents you embedded elsewhere into an MVS namespace and search them by raw vector through the features search endpoint. Hybrid BM25 plus dense is not part of a plain BYO upsert: the documents carry dense vectors only and no text index is created. If you want a lexical leg later, declare a TEXT payload index on the field when you create the namespace; this recipe shows that declaration and the dense search that works today.

    text
    Single Tier

    "FastAPI Pydantic v2 validation patterns"

    Why This Matters

    Pure vector search misses exact identifiers and error codes, and teams that bring their own embeddings often assume keyword matching comes with the vector store. On BYO documents it does not unless the text index exists before the first upsert. Declaring it up front is cheap; discovering its absence after indexing a corpus is not.

    import requests
    from openai import OpenAI
    from mixpeek import Mixpeek
    openai = OpenAI(api_key="YOUR_OPENAI_KEY")
    client = Mixpeek(api_key="YOUR_API_KEY")
    NAMESPACE = "byo-hybrid"
    HEADERS = {"Authorization": "Bearer YOUR_API_KEY"}
    def embed(text):
    return openai.embeddings.create(model="text-embedding-3-small", input=text).data[0].embedding
    # 1. A standalone namespace for your vectors
    client.namespaces.create(
    namespace_id=NAMESPACE,
    mode="standalone",
    vector_configs=[{"name": "dense", "dimension": 1536, "metric": "cosine"}],
    )
    # 2. BM25 reads the namespace's text payload indexes, so declare one on the field that
    # holds the text before the first upsert. The SDK has no namespace update, so this is REST.
    requests.patch(
    f"https://api.mixpeek.com/v1/namespaces/{NAMESPACE}",
    headers=HEADERS,
    json={"payload_indexes": [{"field_name": "content", "type": "text"}]},
    )
    # 3. Upsert vectors with the text in the indexed payload field
    documents = [
    "FastAPI uses Pydantic v2 for data validation and serialization",
    "Express.js middleware handles request and response transformations",
    "Django ORM provides database abstraction with the QuerySet API",
    ]
    client.namespaces.documents.upsert(
    namespace_id=NAMESPACE,
    documents=[
    {"document_id": f"doc-{i}", "vectors": {"dense": embed(text)}, "payload": {"content": text}}
    for i, text in enumerate(documents)
    ],
    )
    # 4. One stage runs both legs: dense on the vector, lexical on the text, fused with RRF
    retriever = client.retrievers.create(
    retriever_name="byo-hybrid-search",
    input_schema={"qv": {"type": "array", "required": True}, "query": {"type": "text", "required": True}},
    stages=[
    {
    "stage_name": "search",
    "stage_id": "feature_search",
    "parameters": {
    "searches": [
    {
    "feature_uri": "dense",
    "query": {"input_mode": "vector", "value": "{{INPUT.qv}}"},
    "top_k": 20,
    },
    {
    "feature_uri": "dense",
    "query": {"input_mode": "text", "value": "{{INPUT.query}}"},
    "top_k": 20,
    "lexical": True,
    },
    ],
    "fusion": "rrf",
    "final_top_k": 5,
    },
    },
    ],
    )
    query = "FastAPI Pydantic validation"
    results = client.retrievers.execute(retriever["retriever_id"], inputs={"qv": embed(query), "query": query})
    for doc in results["documents"]:
    print(round(doc["score"], 3), doc.get("content"))

    Feature Extractors

    Retriever Stages

    feature search

    Search and filter documents by vector similarity using feature embeddings

    filter