pplx-embed-v2-late-0.6b
by perplexity-ai
A 0.6B late-interaction retriever for text, images and page screenshots, 62.3 nDCG@10 on ViDoRe v3 images
perplexity-ai/pplx-embed-v2-late-0.6bDeploy pplx-embed-v2-late-0.6b
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
pplx-embed-v2-late-0.6b is a multimodal late-interaction (ColBERT-style) retriever from Perplexity, built on Qwen3.5 with bidirectional attention and released under the MIT license. It turns a query or a document into one 128-dimensional vector per token and scores a match with MaxSim, so a query can match the specific words or page regions it is about.\n\nIt embeds text, images and visual documents such as page screenshots, but not text and an image in the same input. It shares an embedding space with the larger pplx-embed-v2-late-9b, so this 0.6B model can query an index built with the 9B. The card reports 62.3% nDCG@10 on public ViDoRe v3 with page images and 61.2% with markdown.
Architecture
Qwen3.5 backbone with bidirectional attention, 340M active parameters. Outputs one 128-dimensional vector per token and scores query against document with MaxSim. Distilled from an internal 18B ColBERT teacher with a token-level LEAF-style objective; the 0.6B model was fully fine-tuned.
Mixpeek SDK Integration
# No Mixpeek extractor takes a Hugging Face model id, so this model runs on your side.
# It returns one 128-d vector per token; a single-vector index needs them pooled,
# which gives up the token-level matching (see the FAQ below).
from PIL import Image
from sentence_transformers import MultiVectorEncoder
import requests
model = MultiVectorEncoder("perplexity-ai/pplx-embed-v2-late-0.6b", device="cuda")
page = Image.open("report-2026-q2-page-14.png").convert("RGB")
token_vectors = model.encode_document([page])[0] # (num_tokens, 128)
requests.post(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
headers={"Authorization": "Bearer API_KEY"},
json={"collection_id": "col_your_collection",
"documents": [{"document_id": "report-2026-q2-page-14",
"vectors": {"page-embedding": token_vectors.mean(axis=0).tolist()},
"payload": {"source_key": "reports/2026-q2.pdf#page=14"}}]},
)Capabilities
- One 128-d vector per token, scored with MaxSim (late interaction)
- Text, image and visual-document inputs, encoded in separate batches
- Shares an embedding space with pplx-embed-v2-late-9b
- Native Sentence Transformers MultiVectorEncoder, no custom code
- MIT license
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Public ViDoRe v3, image | nDCG@10 | 62.3% | Model card: perplexity-ai/pplx-embed-v2-late-0.6b (self-reported; 9B model 65.2%) |
| Public ViDoRe v3, markdown | nDCG@10 | 61.2% | Model card (self-reported; 9B model 64.7%) |
Performance
Requires sentence-transformers 6.0.0 or later and transformers 5.4.0 or later. The card notes that PyLate inserts its query and document markers in the second position while this model expects them first. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
What is pplx-embed-v2-late-0.6b?
A late-interaction retrieval model from Perplexity that embeds text, images and page screenshots as one 128-dimensional vector per token and scores matches with MaxSim. It is MIT licensed.
How does it score on visual document retrieval?
The card reports 62.3% nDCG@10 on public ViDoRe v3 using page images and 61.2% using markdown. The 9B sibling reports 65.2% and 64.7%.
Can I query a 9B index with the 0.6B model?
Yes. The card says the 0.6B and 9B models share an embedding space, so the small model can encode queries against an index built with the large one.
Can I store its vectors in a single-vector index?
Only by pooling the per-token vectors into one, which gives up the token-level matching that late interaction is for. The common pattern is a cheaper single-vector first stage, with a late-interaction model rescoring the top candidates.
Specification
Research Paper
Multimodal embeddings beyond a single vector (Perplexity)
arxiv.orgBuild a pipeline with pplx-embed-v2-late-0.6b
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free