Ovis-VL-Embedding-9B
by ATH-MaaS
The 9B Ovis embedding model: text, images, documents and video in one 4096-dimension space
ATH-MaaS/Ovis-VL-Embedding-9BDeploy Ovis-VL-Embedding-9B
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Ovis-VL-Embedding-9B maps text, images, visual documents and video into one vector space, so a written question can retrieve a photo, a PDF page or a clip from a single index. It is the larger sibling of Ovis-VL-Embedding-2B, built on Qwen3.5-9B, with 4096-dimension vectors ranked by cosine similarity. ATH-MaaS released it on 21 September 2026 under Apache-2.0.
The card reports 81.13 overall on MMEB-v2, a 78-dataset image, video and visual-document benchmark, 1.04 points above the strongest baseline it compares against and 3.67 above the 2B model. It leads on images and documents and trails the best baseline on video by 3.05 points.
It does not take audio, the vectors are twice the size of the 2B model's, and the scores are self-reported.
Architecture
Initialized from Qwen3.5-9B with the text and vision encoders and the shared multimodal backbone kept, and the language-modeling head removed. The backbone has 32 layers at hidden size 4096, repeating three Gated DeltaNet layers and one gated full-attention layer. Inputs of any supported type go in as one interleaved sequence, and the embedding is the final-layer hidden state at the last non-padding token, L2-normalized, with no projection head. Training follows the same three stages as the 2B model: multimodal contrastive pretraining, full-parameter finetuning on single-dataset batches, and embedding distillation.
Mixpeek SDK Integration
# No Mixpeek extractor runs these weights. Encode each item as the card describes
# (native processor and chat template, last non-padding token, L2-normalized, 4096
# floats) and upsert into a namespace whose vector index is declared at 4096.
import requests
requests.post(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
headers={"Authorization": "Bearer API_KEY"},
json={
"collection_id": "col_your_collection",
"documents": [{
"document_id": "page-0042",
"vectors": {"ovis_vl_9b": vector}, # the name of your 4096-d index
"payload": {"source_key": "s3://docs/report.pdf", "page": 42},
}],
},
)Capabilities
- One vector space for text, images, visual documents, video frames and interleaved inputs
- Bi-encoder retrieval ranked by cosine similarity
- 4096-dimension output with no projection head
- Apache-2.0 license
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| MMEB-v2 (78 datasets) | Overall | 81.13 | Model card: ATH-MaaS/Ovis-VL-Embedding-9B (self-reported) |
| MMEB-v2 image | Group score | 83.96 | Model card (self-reported) |
| MMEB-v2 visual document | Group score | 83.06 | Model card (self-reported) |
| MMEB-v2 video | Group score | 72.90 | Model card (self-reported; below its best baseline at 75.95) |
Performance
We have not measured latency or memory. At 4096 float32 dimensions each vector is 16 KiB, twice the 2B model's 8 KiB.
Common Pipeline Companions
Frequently Asked Questions
Should I use Ovis-VL-Embedding-9B or the 2B version?
The 9B reports 81.13 overall on MMEB-v2 against 77.46 for the 2B, and its vectors are 4096 dimensions against 2048, so twice the storage per item and more compute to encode. Use the 9B when retrieval quality matters more than index size, and test both on your own content before deciding.
Does Ovis-VL-Embedding-9B handle video and audio?
Video, as sampled frames, yes; audio, no. On video the card reports it 3.05 points behind the strongest baseline it compares against, while it leads on images and visual documents. For audio the card points to Ovis-Omni-Embedding-3B.
Does Mixpeek run Ovis-VL-Embedding-9B?
Not on the managed tier. Encode with it yourself, store one 4096-dimension vector per item in a Mixpeek namespace, and search with a retriever that takes a query vector. On a single-tenant Enterprise deployment the weights can be uploaded and run by a custom plugin.
Specification
Research Paper
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings (arXiv 2609.25165)
arxiv.orgBuild a pipeline with Ovis-VL-Embedding-9B
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free