NEWVectors or files. Pick a path.Start →
    Models/perplexity-ai/pplx-embed-v2-context-9b-preview
    MIT

    pplx-embed-v2-context-9b-preview

    by perplexity-ai

    Contextual chunk embeddings: each chunk is encoded with the rest of its document, at 2048 or 1024 dimensions

    Identifiers
    Model ID
    perplexity-ai/pplx-embed-v2-context-9b-preview
    Feature URI

    Deploy pplx-embed-v2-context-9b-preview

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    pplx-embed-v2-context-9b-preview embeds the chunks of a document together, so each chunk's vector reflects the text around it. That helps retrieval when a chunk on its own is ambiguous, which is most chunks in long reports and contracts. Perplexity released it on 25 September 2026 under the MIT license as a preview.

    It outputs 2048 dimensions, or 1024 by truncation, as unnormalized int8 values, and encodes queries with a separate method from documents.

    The card reports no benchmark scores, and as a preview its weights and embeddings may change without backward compatibility.

    Architecture

    A roughly 8.4B-parameter transformer with custom code (trust_remote_code). A document is passed as a list of chunks and encoded in one pass, and each chunk's embedding is the mean of its token states, so it reflects the whole document. Queries use fixed query prefixes through encode_queries. Trained with Matryoshka losses at 1024 and 2048 dimensions and quantized to int8.

    Mixpeek SDK Integration

    # No Mixpeek extractor takes a Hugging Face model id, so this model runs on your side
    # and its chunk vectors are upserted into a bring-your-own-vectors collection.
    from transformers import AutoModel
    import requests
    
    model = AutoModel.from_pretrained("perplexity-ai/pplx-embed-v2-context-9b-preview", trust_remote_code=True).to("cuda")
    chunks = ["Q3 revenue grew 14 percent.", "Most of the growth came from EMEA.", "Churn fell to 2 percent."]
    vectors = model.encode([chunks], normalize_embeddings=True)[0]  # one 2048-d vector per chunk, in context
    
    requests.post(
        "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
        headers={"Authorization": "Bearer API_KEY"},
        json={"collection_id": "col_your_collection",
              "documents": [{"document_id": f"report-q3-{i}", "vectors": {"text-embedding": v.tolist()},
                             "payload": {"text": c, "chunk": i}} for i, (c, v) in enumerate(zip(chunks, vectors))]},
    )

    Capabilities

    • Contextual embeddings for document chunks, encoded together per document
    • Separate query encoding (encode_queries) from document encoding (encode)
    • 2048 dimensions, or 1024 by truncation (Matryoshka); int8 output
    • Multilingual; MIT license; preview release

    Use Cases on Mixpeek

    Retrieval over long documents where a chunk alone is ambiguous ("it", "the company", "this quarter")
    RAG over reports, contracts and transcripts split into chunks
    Testing contextual embeddings against a plain chunk embedding on your own corpus

    Performance

    Input SizeDocuments as lists of text chunks, encoded together; queries encoded separately
    Embedding Dim2048 (1024 by Matryoshka truncation)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    The card reports no benchmark scores for this preview. Output is unnormalized int8; compare with cosine similarity or normalize first. Preview embeddings should not be mixed with a later version's. Needs a GPU for practical throughput. We have not measured it.

    Frequently Asked Questions

    What is a contextual embedding model?

    It embeds each chunk of a document with the other chunks in view, so a chunk that says "revenue grew 14 percent" carries which company and which quarter from elsewhere in the document. A plain embedding model sees each chunk alone.

    How do I encode queries with pplx-embed-v2-context-9b-preview?

    Use encode_queries for queries and encode for document chunks. The model was trained with different prefixes for each, and the card warns that encoding queries with encode gives worse results.

    How many dimensions does it output?

    2048, or 1024 if you take the first 1024 values and normalize afterwards. Those are the two sizes it was trained for; other truncations were not trained.

    Can I use it in production?

    It is a preview. Perplexity says weights, embeddings and the interface may change without backward compatibility, so do not mix its vectors with a later version's, and plan to re-embed when the final model ships.

    Specification

    Organizationperplexity-ai
    Retriever-
    Parameters8.4B
    LicenseMIT
    Downloads/moN/A
    Likes52

    Research Paper

    pplx-embed-v2-context-9b-preview model card

    arxiv.org

    Build a pipeline with pplx-embed-v2-context-9b-preview

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free