NEWVectors or files. Pick a path.Start →
    Cross-MediaSimilar

    RAG with MVS Standalone

    Complete RAG pipeline using MVS for retrieval and OpenAI for generation. Chunk your documents, embed them with any provider, store in MVS, retrieve relevant context, and generate answers -- no managed feature extractors needed.

    text
    Multi-Tier

    "What is the recommended database architecture for high availability?"

    Why This Matters

    Full control over your RAG pipeline without vendor lock-in. Choose your own chunking strategy, embedding model, and LLM while MVS handles the vector storage and retrieval at scale.

    from openai import OpenAI
    from mixpeek import Mixpeek
    openai = OpenAI(api_key="YOUR_OPENAI_KEY")
    client = Mixpeek(api_key="YOUR_API_KEY")
    NAMESPACE = "rag-docs"
    def embed(text):
    return openai.embeddings.create(model="text-embedding-3-small", input=text).data[0].embedding
    def chunk_text(text, size=200, overlap=40):
    words = text.split()
    return [" ".join(words[i:i + size]) for i in range(0, len(words), size - overlap)]
    client.namespaces.create(
    namespace_id=NAMESPACE,
    mode="standalone",
    vector_configs=[{"name": "dense", "dimension": 1536, "metric": "cosine"}],
    )
    # 1. Chunk a document and upsert every chunk with its vector and text
    with open("architecture-guide.md") as f:
    chunks = chunk_text(f.read())
    client.namespaces.documents.upsert(
    namespace_id=NAMESPACE,
    documents=[
    {
    "document_id": f"architecture-guide-{i}",
    "vectors": {"dense": embed(chunk)},
    "payload": {"content": chunk, "source": "architecture-guide.md", "chunk_index": i},
    }
    for i, chunk in enumerate(chunks)
    ],
    )
    # 2. Retrieve the closest chunks for the question
    question = "What is the recommended database architecture for high availability?"
    results = client.search(
    namespace_id=NAMESPACE,
    queries=[{"feature_uri": "dense", "query": {"input_mode": "vector", "value": embed(question)}, "top_k": 5}],
    )
    # 3. Answer from those chunks, citing them by index
    context = " ".join(f"[{doc.get('chunk_index')}] {doc.get('content')}" for doc in results["documents"])
    response = openai.chat.completions.create(
    model="gpt-4o",
    messages=[
    {"role": "system", "content": "Answer from this context and cite chunk numbers. " + context},
    {"role": "user", "content": question},
    ],
    )
    print(response.choices[0].message.content)

    Feature Extractors

    Retriever Stages

    feature search

    Search and filter documents by vector similarity using feature embeddings

    filter

    Related Recipes & Resources

    Explore these related resources to deepen your understanding and discover more powerful features