NEWVectors or files. Pick a path.Start →
    Models/perplexity-ai/pplx-embed-v2-late-9b
    MIT

    pplx-embed-v2-late-9b

    by perplexity-ai

    Perplexity's 9B late-interaction retriever for text, images and page screenshots, 65.2 nDCG@10 on ViDoRe v3 images

    30likes
    9B (7.4B active)params
    Identifiers
    Model ID
    perplexity-ai/pplx-embed-v2-late-9b
    Feature URI

    Deploy pplx-embed-v2-late-9b

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    pplx-embed-v2-late-9b is the larger of Perplexity's two multimodal late-interaction (ColBERT-style) retrievers, built on Qwen3.5 with bidirectional attention and released under the MIT license. It turns a query or a document into one 128-dimensional vector per token and scores a match with MaxSim, so a query can match the specific words or page regions it is about.\n\nIt embeds text, images and visual documents such as page screenshots, but not text and an image in the same input. The card reports 65.2% nDCG@10 on public ViDoRe v3 with page images and 64.7% with markdown, about three points above the 0.6B model, which shares its embedding space.

    Architecture

    Qwen3.5 backbone with bidirectional attention, 7.4B active parameters (8.4B in the checkpoint). Outputs one 128-dimensional vector per token and scores query against document with MaxSim. Distilled from an internal 18B ColBERT teacher with a token-level LEAF-style objective; for this model the final eight transformer layers were fully trained.

    Mixpeek SDK Integration

    # No Mixpeek extractor takes a Hugging Face model id, so this model runs on your side.
    # It returns one 128-d vector per token; a single-vector index needs them pooled,
    # which gives up the token-level matching (see the FAQ below).
    from PIL import Image
    from sentence_transformers import MultiVectorEncoder
    import requests
    
    model = MultiVectorEncoder("perplexity-ai/pplx-embed-v2-late-9b", device="cuda")
    page = Image.open("report-2026-q2-page-14.png").convert("RGB")
    token_vectors = model.encode_document([page])[0]  # (num_tokens, 128)
    
    requests.post(
        "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
        headers={"Authorization": "Bearer API_KEY"},
        json={"collection_id": "col_your_collection",
              "documents": [{"document_id": "report-2026-q2-page-14",
                             "vectors": {"page-embedding": token_vectors.mean(axis=0).tolist()},
                             "payload": {"source_key": "reports/2026-q2.pdf#page=14"}}]},
    )

    Capabilities

    • One 128-d vector per token, scored with MaxSim (late interaction)
    • Text, image and visual-document inputs, encoded in separate batches
    • Shares an embedding space with pplx-embed-v2-late-0.6b
    • Native Sentence Transformers MultiVectorEncoder, no custom code
    • MIT license

    Use Cases on Mixpeek

    Indexing scanned reports, slides and PDFs as page images at the family's best accuracy
    Building the index with the 9B and encoding queries with the faster 0.6B
    Reranking the top candidates from a cheaper single-vector search

    Benchmarks

    DatasetMetricScoreSource
    Public ViDoRe v3, imagenDCG@1065.2%Model card: perplexity-ai/pplx-embed-v2-late-9b (self-reported; 0.6B model 62.3%)
    Public ViDoRe v3, markdownnDCG@1064.7%Model card (self-reported; 0.6B model 61.2%)

    Performance

    Input SizeText, images or page screenshots; text-only and image-only batches must be encoded separately
    Embedding Dim128 per token (multi-vector)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    Requires sentence-transformers 6.0.0 or later and transformers 5.4.0 or later. Token-level vectors take far more storage than one vector per page. We have not measured it.

    Frequently Asked Questions

    What is pplx-embed-v2-late-9b?

    Perplexity's larger late-interaction retrieval model. It embeds text, images and page screenshots as one 128-dimensional vector per token, scores matches with MaxSim, and is MIT licensed.

    How much better is the 9B than the 0.6B?

    On public ViDoRe v3 the card reports 65.2% nDCG@10 with page images and 64.7% with markdown for the 9B, against 62.3% and 61.2% for the 0.6B, about three points, for roughly 20 times the active parameters.

    Can I build the index with the 9B and query with the 0.6B?

    Yes. The card says the two models share an embedding space, so you can embed documents once with the 9B and encode queries with the smaller, faster 0.6B.

    Can I store its vectors in a single-vector index?

    Only by pooling the per-token vectors into one, for example by averaging, which gives up the token-level matching that makes late interaction accurate. Use the pooled vector for a first pass and rescore the top results with MaxSim over the full token vectors.

    Specification

    Organizationperplexity-ai
    Retriever-
    Parameters9B (7.4B active)
    LicenseMIT
    Downloads/moN/A
    Likes30

    Research Paper

    Multimodal embeddings beyond a single vector (Perplexity)

    arxiv.org

    Build a pipeline with pplx-embed-v2-late-9b

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free