NEWVectors or files. Pick a path.Start →
    Models/Embeddings/jinaai/jina-embeddings-v5-omni-nano
    HFVisual Embeddingscc-by-nc-4.0

    jina-embeddings-v5-omni-nano

    by jinaai

    Compact omni-modal embedding model for text, images, video, and audio in one vector space

    8Kdl/month
    40likes
    986Mparams
    Identifiers
    Model ID
    jinaai/jina-embeddings-v5-omni-nano
    Feature URI
    mixpeek://image_extractor@v1/jina_embeddings_v5_omni_nano

    Deploy jina-embeddings-v5-omni-nano

    Single-tenant

    Mixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.

    Overview

    Jina Embeddings v5 Omni Nano is the smallest model in the Jina v5 omni family, placing text, images, video frames, and audio into a single shared vector space. At ~239M parameters, it runs efficiently on edge devices and high-throughput pipelines.

    The model shares the same text embedding space as jina-v5-text, meaning existing text indexes remain backwards-compatible when adding multimodal content. This makes it the lowest-friction path to cross-modal search.

    Architecture

    Multimodal transformer encoder with separate input projections for text, image, video, and audio modalities. All modalities project into a shared embedding space. Matryoshka representation learning enables flexible output dimensions.

    Mixpeek SDK Integration

    // No extractor parameter takes a Hugging Face model id (checked against
    // GET /v1/discovery/extractors, which returns 13), so jina-embeddings-v5-omni-nano runs
    // on your side and the output is upserted through POST
    // /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
    // path is to upload the weights instead: POST /v1/namespaces/{id}/models
    // accepts the huggingface format and a custom plugin loads them.
    const res = await fetch(
      "https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
      {
        method: "POST",
        headers: {
          Authorization: "Bearer API_KEY",
          "Content-Type": "application/json",
        },
        body: JSON.stringify({
          collection_id: "col_your_collection",
          documents: [
            {
              document_id: "asset-00412",
              // The vector name has to match a vector index on the collection.
              vectors: { "image-embedding": yourVector },
              payload: { source_key: "archive/2026/asset-00412" },
            },
          ],
        }),
      },
    );
    
    // Managed alternative, if this exact model is not the requirement:
    // image_extractor@v1 runs google/siglip-base-patch16-224
    // (768-d) over a bucket, with no inference of your own.

    Capabilities

    • Omni-modal: text, images, video, audio in one space
    • Backwards-compatible with jina-v5-text indexes
    • ~239M parameters for edge/high-throughput deployment
    • Matryoshka dimensions for flexible storage
    • Apache 2.0 license

    Use Cases on Mixpeek

    Cross-modal search (find images matching text queries, or vice versa)
    High-throughput multimodal indexing where latency matters
    Edge deployment for on-device multimodal understanding

    Benchmarks

    DatasetMetricScoreSource
    Cross-modal retrievalRecall@10Competitive with 677M variantJina AI, May 2026

    Performance

    Input SizeText: 8192 tokens; Image: 224x224+; Audio: 30s clips
    GPU Latency~3ms / item (A100)
    GPU Throughput~3000 items/sec (A100, batch 128)
    GPU Memory~0.5 GB

    Specification

    FrameworkHF
    Organizationjinaai
    FeatureVisual Embeddings
    Output768-dim vector
    Modalitiesvideo, image
    RetrieverVector Search
    Parameters986M
    Licensecc-by-nc-4.0
    Downloads/mo8K
    Likes40

    Research Paper

    Jina Embeddings v5 Omni: Multimodal Embeddings for Text, Image, Audio, and Video

    arxiv.org

    Build a pipeline with jina-embeddings-v5-omni-nano

    Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.

    Run it on your own data, free