Ovis-VL-Embedding-2B
by ATH-MaaS
A 2B embedding model that puts text, images, documents and video in one vector space
ATH-MaaS/Ovis-VL-Embedding-2BDeploy Ovis-VL-Embedding-2B
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Ovis-VL-Embedding-2B turns text, images, visual documents and video into vectors in one shared space, so a written question can find a photo, a PDF page or a video clip with a single index. It is a compact bi-encoder: queries and candidates are encoded separately, each becomes one 2048-dimension vector, and results are ranked by cosine similarity. ATH-MaaS released it on 21 September 2026 under Apache-2.0, alongside a 9B version and the audio-capable Ovis-Omni-Embedding-3B.
The card reports 77.46 overall on MMEB-v2, a 78-dataset benchmark across image, video and visual-document tasks, 2.04 points above the strongest baseline it compares against. Its lead is widest on images (+3.21); on video it sits 1.72 points behind the best baseline.
It does not handle audio, and the scores are self-reported by the authors. Test it on your own content before relying on the numbers.
Architecture
Initialized from Qwen3.5-2B, keeping its text and vision encoders and the shared multimodal backbone, with the language-modeling head removed. The backbone has 24 layers at hidden size 2048, repeating three Gated DeltaNet layers and one gated full-attention layer. Text, images, document pages and sampled video frames go in as one interleaved sequence, and the embedding is the final-layer hidden state at the last non-padding token, L2-normalized, with no modality-specific projection head. Training ran in three stages: multimodal contrastive pretraining, full-parameter finetuning on single-dataset batches, and embedding distillation from expert models.
Mixpeek SDK Integration
# No Mixpeek extractor runs these weights. Encode each item with Ovis-VL-Embedding-2B
# as its card describes (native processor and chat template, last non-padding token
# of the final layer, L2-normalized, 2048 floats), then upsert into a namespace whose
# vector index is declared at 2048 dimensions.
import requests
requests.post(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
headers={"Authorization": "Bearer API_KEY"},
json={
"collection_id": "col_your_collection",
"documents": [{
"document_id": "photo-0042",
"vectors": {"ovis_vl": vector}, # the name of your 2048-d index
"payload": {"source_key": "s3://media/photo-0042.jpg"},
}],
},
)Capabilities
- One vector space for text, images, visual documents, video frames and interleaved inputs
- Bi-encoder retrieval: encode once, rank by cosine similarity
- 2048-dimension output with no projection head
- Apache-2.0 license
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| MMEB-v2 (78 datasets) | Overall | 77.46 | Model card: ATH-MaaS/Ovis-VL-Embedding-2B (self-reported) |
| MMEB-v2 image | Group score | 80.62 | Model card (self-reported) |
| MMEB-v2 visual document | Group score | 80.47 | Model card (self-reported) |
| MMEB-v2 video | Group score | 67.12 | Model card (self-reported; below its best baseline at 68.84) |
Performance
We have not measured encoding latency or memory. At 2048 float32 dimensions each vector is 8 KiB before any compression.
Common Pipeline Companions
Frequently Asked Questions
What can Ovis-VL-Embedding-2B search?
Text, images, visual documents such as PDF pages and slides, and video represented by sampled frames, all in one vector space, so a text query can return an image, a page or a clip from the same index. It does not take audio; the card points to Ovis-Omni-Embedding-3B for that.
What is the difference between Ovis-VL-Embedding-2B and Ovis-Omni-Embedding-3B?
The 2B model is built on Qwen3.5-2B and covers text, images, documents and video. The Omni 3B model is built on Qwen2.5-Omni-3B and adds audio. Both output 2048-dimension vectors and are ranked by cosine similarity, so the choice mostly depends on whether audio needs to share the index.
Are its benchmark results independently verified?
No. The MMEB-v2 scores are reported on the model card, which also notes that video question answering, general video retrieval and ViDoRe-V2 remain below the strongest specialist baselines. MMEB-v2 is public, so the numbers can be reproduced; until they are, test on a sample of your own content.
Does Mixpeek run Ovis-VL-Embedding-2B?
Not on the managed tier. Encode with it yourself, store one 2048-dimension vector per item in a Mixpeek namespace, and search with a retriever that takes a query vector. On a single-tenant Enterprise deployment the weights can be uploaded and run by a custom plugin.
Specification
Research Paper
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings (arXiv 2609.25165)
arxiv.orgBuild a pipeline with Ovis-VL-Embedding-2B
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free