NEWVectors or files. Pick a path.Start →
    Models/ATH-MaaS/Ovis-VL-Embedding-2B
    Apache-2.0

    Ovis-VL-Embedding-2B

    by ATH-MaaS

    A 2B embedding model that puts text, images, documents and video in one vector space

    Identifiers
    Model ID
    ATH-MaaS/Ovis-VL-Embedding-2B
    Feature URI

    Deploy Ovis-VL-Embedding-2B

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Ovis-VL-Embedding-2B turns text, images, visual documents and video into vectors in one shared space, so a written question can find a photo, a PDF page or a video clip with a single index. It is a compact bi-encoder: queries and candidates are encoded separately, each becomes one 2048-dimension vector, and results are ranked by cosine similarity. ATH-MaaS released it on 21 September 2026 under Apache-2.0, alongside a 9B version and the audio-capable Ovis-Omni-Embedding-3B.

    The card reports 77.46 overall on MMEB-v2, a 78-dataset benchmark across image, video and visual-document tasks, 2.04 points above the strongest baseline it compares against. Its lead is widest on images (+3.21); on video it sits 1.72 points behind the best baseline.

    It does not handle audio, and the scores are self-reported by the authors. Test it on your own content before relying on the numbers.

    Architecture

    Initialized from Qwen3.5-2B, keeping its text and vision encoders and the shared multimodal backbone, with the language-modeling head removed. The backbone has 24 layers at hidden size 2048, repeating three Gated DeltaNet layers and one gated full-attention layer. Text, images, document pages and sampled video frames go in as one interleaved sequence, and the embedding is the final-layer hidden state at the last non-padding token, L2-normalized, with no modality-specific projection head. Training ran in three stages: multimodal contrastive pretraining, full-parameter finetuning on single-dataset batches, and embedding distillation from expert models.

    Mixpeek SDK Integration

    # No Mixpeek extractor runs these weights. Encode each item with Ovis-VL-Embedding-2B
    # as its card describes (native processor and chat template, last non-padding token
    # of the final layer, L2-normalized, 2048 floats), then upsert into a namespace whose
    # vector index is declared at 2048 dimensions.
    import requests
    
    requests.post(
        "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
        headers={"Authorization": "Bearer API_KEY"},
        json={
            "collection_id": "col_your_collection",
            "documents": [{
                "document_id": "photo-0042",
                "vectors": {"ovis_vl": vector},  # the name of your 2048-d index
                "payload": {"source_key": "s3://media/photo-0042.jpg"},
            }],
        },
    )

    Capabilities

    • One vector space for text, images, visual documents, video frames and interleaved inputs
    • Bi-encoder retrieval: encode once, rank by cosine similarity
    • 2048-dimension output with no projection head
    • Apache-2.0 license

    Use Cases on Mixpeek

    Text-to-image and image-to-image search over a photo or product library
    Finding a PDF page or slide from a question, without an OCR step
    Searching short video by description, alongside images and documents in one index
    Multimodal RAG where one encoder covers every non-audio file type

    Benchmarks

    DatasetMetricScoreSource
    MMEB-v2 (78 datasets)Overall77.46Model card: ATH-MaaS/Ovis-VL-Embedding-2B (self-reported)
    MMEB-v2 imageGroup score80.62Model card (self-reported)
    MMEB-v2 visual documentGroup score80.47Model card (self-reported)
    MMEB-v2 videoGroup score67.12Model card (self-reported; below its best baseline at 68.84)

    Performance

    Input SizeText, images, visual documents, sampled video frames and interleaved inputs; no audio
    Embedding Dim2048
    GPU LatencyInput dependent
    GPU ThroughputBatch dependent
    GPU MemoryModel dependent

    We have not measured encoding latency or memory. At 2048 float32 dimensions each vector is 8 KiB before any compression.

    Frequently Asked Questions

    What can Ovis-VL-Embedding-2B search?

    Text, images, visual documents such as PDF pages and slides, and video represented by sampled frames, all in one vector space, so a text query can return an image, a page or a clip from the same index. It does not take audio; the card points to Ovis-Omni-Embedding-3B for that.

    What is the difference between Ovis-VL-Embedding-2B and Ovis-Omni-Embedding-3B?

    The 2B model is built on Qwen3.5-2B and covers text, images, documents and video. The Omni 3B model is built on Qwen2.5-Omni-3B and adds audio. Both output 2048-dimension vectors and are ranked by cosine similarity, so the choice mostly depends on whether audio needs to share the index.

    Are its benchmark results independently verified?

    No. The MMEB-v2 scores are reported on the model card, which also notes that video question answering, general video retrieval and ViDoRe-V2 remain below the strongest specialist baselines. MMEB-v2 is public, so the numbers can be reproduced; until they are, test on a sample of your own content.

    Does Mixpeek run Ovis-VL-Embedding-2B?

    Not on the managed tier. Encode with it yourself, store one 2048-dimension vector per item in a Mixpeek namespace, and search with a retriever that takes a query vector. On a single-tenant Enterprise deployment the weights can be uploaded and run by a custom plugin.

    Specification

    OrganizationATH-MaaS
    Retriever-
    Parameters2B
    LicenseApache-2.0
    Downloads/moN/A
    Likes33

    Research Paper

    Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings (arXiv 2609.25165)

    arxiv.org

    Build a pipeline with Ovis-VL-Embedding-2B

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free