NEWVectors or files. Pick a path.Start →
    Models/ATH-MaaS/Ovis-VL-Embedding-9B
    Apache-2.0

    Ovis-VL-Embedding-9B

    by ATH-MaaS

    The 9B Ovis embedding model: text, images, documents and video in one 4096-dimension space

    Identifiers
    Model ID
    ATH-MaaS/Ovis-VL-Embedding-9B
    Feature URI

    Deploy Ovis-VL-Embedding-9B

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Ovis-VL-Embedding-9B maps text, images, visual documents and video into one vector space, so a written question can retrieve a photo, a PDF page or a clip from a single index. It is the larger sibling of Ovis-VL-Embedding-2B, built on Qwen3.5-9B, with 4096-dimension vectors ranked by cosine similarity. ATH-MaaS released it on 21 September 2026 under Apache-2.0.

    The card reports 81.13 overall on MMEB-v2, a 78-dataset image, video and visual-document benchmark, 1.04 points above the strongest baseline it compares against and 3.67 above the 2B model. It leads on images and documents and trails the best baseline on video by 3.05 points.

    It does not take audio, the vectors are twice the size of the 2B model's, and the scores are self-reported.

    Architecture

    Initialized from Qwen3.5-9B with the text and vision encoders and the shared multimodal backbone kept, and the language-modeling head removed. The backbone has 32 layers at hidden size 4096, repeating three Gated DeltaNet layers and one gated full-attention layer. Inputs of any supported type go in as one interleaved sequence, and the embedding is the final-layer hidden state at the last non-padding token, L2-normalized, with no projection head. Training follows the same three stages as the 2B model: multimodal contrastive pretraining, full-parameter finetuning on single-dataset batches, and embedding distillation.

    Mixpeek SDK Integration

    # No Mixpeek extractor runs these weights. Encode each item as the card describes
    # (native processor and chat template, last non-padding token, L2-normalized, 4096
    # floats) and upsert into a namespace whose vector index is declared at 4096.
    import requests
    
    requests.post(
        "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
        headers={"Authorization": "Bearer API_KEY"},
        json={
            "collection_id": "col_your_collection",
            "documents": [{
                "document_id": "page-0042",
                "vectors": {"ovis_vl_9b": vector},  # the name of your 4096-d index
                "payload": {"source_key": "s3://docs/report.pdf", "page": 42},
            }],
        },
    )

    Capabilities

    • One vector space for text, images, visual documents, video frames and interleaved inputs
    • Bi-encoder retrieval ranked by cosine similarity
    • 4096-dimension output with no projection head
    • Apache-2.0 license

    Use Cases on Mixpeek

    Image and document retrieval where accuracy matters more than index size
    Finding a PDF page or slide from a question, without OCR
    Measuring how much a smaller embedding model gives up on your own content
    Multimodal RAG over images, documents and short video in one index

    Benchmarks

    DatasetMetricScoreSource
    MMEB-v2 (78 datasets)Overall81.13Model card: ATH-MaaS/Ovis-VL-Embedding-9B (self-reported)
    MMEB-v2 imageGroup score83.96Model card (self-reported)
    MMEB-v2 visual documentGroup score83.06Model card (self-reported)
    MMEB-v2 videoGroup score72.90Model card (self-reported; below its best baseline at 75.95)

    Performance

    Input SizeText, images, visual documents, sampled video frames and interleaved inputs; no audio
    Embedding Dim4096
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    We have not measured latency or memory. At 4096 float32 dimensions each vector is 16 KiB, twice the 2B model's 8 KiB.

    Frequently Asked Questions

    Should I use Ovis-VL-Embedding-9B or the 2B version?

    The 9B reports 81.13 overall on MMEB-v2 against 77.46 for the 2B, and its vectors are 4096 dimensions against 2048, so twice the storage per item and more compute to encode. Use the 9B when retrieval quality matters more than index size, and test both on your own content before deciding.

    Does Ovis-VL-Embedding-9B handle video and audio?

    Video, as sampled frames, yes; audio, no. On video the card reports it 3.05 points behind the strongest baseline it compares against, while it leads on images and visual documents. For audio the card points to Ovis-Omni-Embedding-3B.

    Does Mixpeek run Ovis-VL-Embedding-9B?

    Not on the managed tier. Encode with it yourself, store one 4096-dimension vector per item in a Mixpeek namespace, and search with a retriever that takes a query vector. On a single-tenant Enterprise deployment the weights can be uploaded and run by a custom plugin.

    Specification

    OrganizationATH-MaaS
    Retriever-
    Parameters9B
    LicenseApache-2.0
    Downloads/moN/A
    Likes35

    Research Paper

    Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings (arXiv 2609.25165)

    arxiv.org

    Build a pipeline with Ovis-VL-Embedding-9B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free