pplx-embed-v2-context-9b-preview
by perplexity-ai
Contextual chunk embeddings: each chunk is encoded with the rest of its document, at 2048 or 1024 dimensions
perplexity-ai/pplx-embed-v2-context-9b-previewDeploy pplx-embed-v2-context-9b-preview
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
pplx-embed-v2-context-9b-preview embeds the chunks of a document together, so each chunk's vector reflects the text around it. That helps retrieval when a chunk on its own is ambiguous, which is most chunks in long reports and contracts. Perplexity released it on 25 September 2026 under the MIT license as a preview.
It outputs 2048 dimensions, or 1024 by truncation, as unnormalized int8 values, and encodes queries with a separate method from documents.
The card reports no benchmark scores, and as a preview its weights and embeddings may change without backward compatibility.
Architecture
A roughly 8.4B-parameter transformer with custom code (trust_remote_code). A document is passed as a list of chunks and encoded in one pass, and each chunk's embedding is the mean of its token states, so it reflects the whole document. Queries use fixed query prefixes through encode_queries. Trained with Matryoshka losses at 1024 and 2048 dimensions and quantized to int8.
Mixpeek SDK Integration
# No Mixpeek extractor takes a Hugging Face model id, so this model runs on your side
# and its chunk vectors are upserted into a bring-your-own-vectors collection.
from transformers import AutoModel
import requests
model = AutoModel.from_pretrained("perplexity-ai/pplx-embed-v2-context-9b-preview", trust_remote_code=True).to("cuda")
chunks = ["Q3 revenue grew 14 percent.", "Most of the growth came from EMEA.", "Churn fell to 2 percent."]
vectors = model.encode([chunks], normalize_embeddings=True)[0] # one 2048-d vector per chunk, in context
requests.post(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
headers={"Authorization": "Bearer API_KEY"},
json={"collection_id": "col_your_collection",
"documents": [{"document_id": f"report-q3-{i}", "vectors": {"text-embedding": v.tolist()},
"payload": {"text": c, "chunk": i}} for i, (c, v) in enumerate(zip(chunks, vectors))]},
)Capabilities
- Contextual embeddings for document chunks, encoded together per document
- Separate query encoding (encode_queries) from document encoding (encode)
- 2048 dimensions, or 1024 by truncation (Matryoshka); int8 output
- Multilingual; MIT license; preview release
Use Cases on Mixpeek
Performance
The card reports no benchmark scores for this preview. Output is unnormalized int8; compare with cosine similarity or normalize first. Preview embeddings should not be mixed with a later version's. Needs a GPU for practical throughput. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
What is a contextual embedding model?
It embeds each chunk of a document with the other chunks in view, so a chunk that says "revenue grew 14 percent" carries which company and which quarter from elsewhere in the document. A plain embedding model sees each chunk alone.
How do I encode queries with pplx-embed-v2-context-9b-preview?
Use encode_queries for queries and encode for document chunks. The model was trained with different prefixes for each, and the card warns that encoding queries with encode gives worse results.
How many dimensions does it output?
2048, or 1024 if you take the first 1024 values and normalize afterwards. Those are the two sizes it was trained for; other truncations were not trained.
Can I use it in production?
It is a preview. Perplexity says weights, embeddings and the interface may change without backward compatibility, so do not mix its vectors with a later version's, and plan to re-embed when the final model ships.
Specification
Research Paper
pplx-embed-v2-context-9b-preview model card
arxiv.orgBuild a pipeline with pplx-embed-v2-context-9b-preview
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free