NEWVectors or files. Pick a path.Start →
    Models/perplexity-ai/pplx-embed-v2-late-0.6b
    MIT

    pplx-embed-v2-late-0.6b

    by perplexity-ai

    A 0.6B late-interaction retriever for text, images and page screenshots, 62.3 nDCG@10 on ViDoRe v3 images

    48likes
    0.6B (340M active)params
    Identifiers
    Model ID
    perplexity-ai/pplx-embed-v2-late-0.6b
    Feature URI

    Deploy pplx-embed-v2-late-0.6b

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    pplx-embed-v2-late-0.6b is a multimodal late-interaction (ColBERT-style) retriever from Perplexity, built on Qwen3.5 with bidirectional attention and released under the MIT license. It turns a query or a document into one 128-dimensional vector per token and scores a match with MaxSim, so a query can match the specific words or page regions it is about.\n\nIt embeds text, images and visual documents such as page screenshots, but not text and an image in the same input. It shares an embedding space with the larger pplx-embed-v2-late-9b, so this 0.6B model can query an index built with the 9B. The card reports 62.3% nDCG@10 on public ViDoRe v3 with page images and 61.2% with markdown.

    Architecture

    Qwen3.5 backbone with bidirectional attention, 340M active parameters. Outputs one 128-dimensional vector per token and scores query against document with MaxSim. Distilled from an internal 18B ColBERT teacher with a token-level LEAF-style objective; the 0.6B model was fully fine-tuned.

    Mixpeek SDK Integration

    # No Mixpeek extractor takes a Hugging Face model id, so this model runs on your side.
    # It returns one 128-d vector per token; a single-vector index needs them pooled,
    # which gives up the token-level matching (see the FAQ below).
    from PIL import Image
    from sentence_transformers import MultiVectorEncoder
    import requests
    
    model = MultiVectorEncoder("perplexity-ai/pplx-embed-v2-late-0.6b", device="cuda")
    page = Image.open("report-2026-q2-page-14.png").convert("RGB")
    token_vectors = model.encode_document([page])[0]  # (num_tokens, 128)
    
    requests.post(
        "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
        headers={"Authorization": "Bearer API_KEY"},
        json={"collection_id": "col_your_collection",
              "documents": [{"document_id": "report-2026-q2-page-14",
                             "vectors": {"page-embedding": token_vectors.mean(axis=0).tolist()},
                             "payload": {"source_key": "reports/2026-q2.pdf#page=14"}}]},
    )

    Capabilities

    • One 128-d vector per token, scored with MaxSim (late interaction)
    • Text, image and visual-document inputs, encoded in separate batches
    • Shares an embedding space with pplx-embed-v2-late-9b
    • Native Sentence Transformers MultiVectorEncoder, no custom code
    • MIT license

    Use Cases on Mixpeek

    Searching scanned reports, slides and PDFs as page images
    Reranking the top candidates from a cheaper single-vector search
    Querying a 9B-built index with a small model at query time

    Benchmarks

    DatasetMetricScoreSource
    Public ViDoRe v3, imagenDCG@1062.3%Model card: perplexity-ai/pplx-embed-v2-late-0.6b (self-reported; 9B model 65.2%)
    Public ViDoRe v3, markdownnDCG@1061.2%Model card (self-reported; 9B model 64.7%)

    Performance

    Input SizeText, images or page screenshots; text-only and image-only batches must be encoded separately
    Embedding Dim128 per token (multi-vector)
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    Requires sentence-transformers 6.0.0 or later and transformers 5.4.0 or later. The card notes that PyLate inserts its query and document markers in the second position while this model expects them first. We have not measured it.

    Frequently Asked Questions

    What is pplx-embed-v2-late-0.6b?

    A late-interaction retrieval model from Perplexity that embeds text, images and page screenshots as one 128-dimensional vector per token and scores matches with MaxSim. It is MIT licensed.

    How does it score on visual document retrieval?

    The card reports 62.3% nDCG@10 on public ViDoRe v3 using page images and 61.2% using markdown. The 9B sibling reports 65.2% and 64.7%.

    Can I query a 9B index with the 0.6B model?

    Yes. The card says the 0.6B and 9B models share an embedding space, so the small model can encode queries against an index built with the large one.

    Can I store its vectors in a single-vector index?

    Only by pooling the per-token vectors into one, which gives up the token-level matching that late interaction is for. The common pattern is a cheaper single-vector first stage, with a late-interaction model rescoring the top candidates.

    Specification

    Organizationperplexity-ai
    Retriever-
    Parameters0.6B (340M active)
    LicenseMIT
    Downloads/moN/A
    Likes48

    Research Paper

    Multimodal embeddings beyond a single vector (Perplexity)

    arxiv.org

    Build a pipeline with pplx-embed-v2-late-0.6b

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free