CLM-v0.1-8B
by Contrastive-LM
Scores and ranks the candidates you give it, returning probabilities instead of generated text
Contrastive-LM/CLM-v0.1-8BDeploy CLM-v0.1-8B
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
CLM-v0.1-8B takes a situation and a set of candidates and says which candidate fits best, with a probability for each. That makes it a reranker for search results, a chooser among an agent's possible next actions, and a classifier for typed questions such as yes/no, pick one of these, or rate on this scale. It does not generate text. Contrastive-LM released it on 21 September 2026 under Apache-2.0.
It is two small projection heads, one for the state and one for the candidate, on top of a frozen Qwen3-8B encoder, trained with a contrastive loss on about 60M question-answer pairs, 30M synthetic hard negatives and 1M agent trajectories. Because states and candidates are encoded separately, candidate embeddings can be cached and reused.
The card's strongest numbers, 81.6% on DeepSWE and 87.6% on Terminal-Bench 2.1, come from heads fine-tuned for those tasks, not from this checkpoint as released.
Architecture
A frozen Qwen3-8B encoder produces last-token pooled embeddings for the state and for each candidate. A state head and an action head project those into a shared space, trained with a bidirectional InfoNCE loss, and a candidate's probability comes from comparing the two projections across the set. Training ran in three stages: pre-training on about 60M Nemotron question-answer pairs, mid-training on about 30M synthetic hard negatives, and post-training on about 1M agentic trajectories. Fine-tuning trains only the heads, which is why it is cheap. The heads only work with Qwen3-8B embeddings.
Mixpeek SDK Integration
# CLM scores candidates at query time, so there is nothing to index. Retrieve with
# Mixpeek first, then rerank the returned documents with CLM on your side.
import requests
from clm import Engine
hits = requests.post(
"https://api.mixpeek.com/v1/retrievers/ret_your_retriever/execute",
headers={"Authorization": "Bearer API_KEY", "X-Namespace": "ns_your_namespace"},
json={"inputs": {"query": question}},
).json()["results"]Capabilities
- Ranks free-form candidates against a question and returns a probability for each
- Typed questions: yes/no, choice from a list, and graded scores
- States and candidates encoded separately, so candidate embeddings can be cached
- Apache-2.0 license; only the small heads are trained when fine-tuning
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| DeepSWE | Verifier accuracy (fine-tuned head) | 81.6% | Model card: Contrastive-LM/CLM-v0.1-8B (self-reported; fine-tuned, not this checkpoint zero-shot) |
| Terminal-Bench 2.1 | Verifier accuracy (fine-tuned head) | 87.6% | Model card (self-reported; fine-tuned) |
Performance
Needs a Qwen3-8B embedding server (the card uses vLLM). Probabilities are relative to the candidate set you pass, so they compare candidates with each other rather than scoring relevance in absolute terms. We have not measured latency.
Common Pipeline Companions
Frequently Asked Questions
What is CLM-v0.1-8B used for?
Scoring and ranking candidates you already have: search results to rerank, tool calls or next moves for an agent to choose between, or labels for a piece of text. It returns a probability per candidate or per answer option and does not generate text.
Is CLM-v0.1-8B a reranker?
It can be used as one. Given a question and a list of candidate passages it returns them ranked with probabilities. Those probabilities are relative to the list you pass, so they tell you which candidate is best in that set and do not measure relevance on an absolute scale.
Are the DeepSWE and Terminal-Bench scores for this checkpoint?
No. The card says those results come from heads fine-tuned on each task, starting from this checkpoint. Zero-shot, the card compares it with other models on computer-use, gaming and tool-calling tasks. The scores are self-reported.
Does Mixpeek run CLM-v0.1-8B?
Not as a built-in stage. Retrieve candidates with a Mixpeek retriever, then rerank them with CLM in your own service, as in the example on this page. Mixpeek retrievers also have their own rerank stage for reranking inside the pipeline.
Specification
Research Paper
Contrastive Language Models: A System One Model for Fast and Generalizable Decision-Making
arxiv.orgBuild a pipeline with CLM-v0.1-8B
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free