jina-embeddings-v5-omni-nano
by jinaai
Compact omni-modal embedding model for text, images, video, and audio in one vector space
jinaai/jina-embeddings-v5-omni-nanomixpeek://image_extractor@v1/jina_embeddings_v5_omni_nanoDeploy jina-embeddings-v5-omni-nano
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
Jina Embeddings v5 Omni Nano is the smallest model in the Jina v5 omni family, placing text, images, video frames, and audio into a single shared vector space. At ~239M parameters, it runs efficiently on edge devices and high-throughput pipelines.
The model shares the same text embedding space as jina-v5-text, meaning existing text indexes remain backwards-compatible when adding multimodal content. This makes it the lowest-friction path to cross-modal search.
Architecture
Multimodal transformer encoder with separate input projections for text, image, video, and audio modalities. All modalities project into a shared embedding space. Matryoshka representation learning enables flexible output dimensions.
Mixpeek SDK Integration
// No extractor parameter takes a Hugging Face model id (checked against
// GET /v1/discovery/extractors, which returns 13), so jina-embeddings-v5-omni-nano runs
// on your side and the output is upserted through POST
// /v1/namespaces/{namespace_id}/documents/upsert. On Enterprise the other
// path is to upload the weights instead: POST /v1/namespaces/{id}/models
// accepts the huggingface format and a custom plugin loads them.
const res = await fetch(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
{
method: "POST",
headers: {
Authorization: "Bearer API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
collection_id: "col_your_collection",
documents: [
{
document_id: "asset-00412",
// The vector name has to match a vector index on the collection.
vectors: { "image-embedding": yourVector },
payload: { source_key: "archive/2026/asset-00412" },
},
],
}),
},
);
// Managed alternative, if this exact model is not the requirement:
// image_extractor@v1 runs google/siglip-base-patch16-224
// (768-d) over a bucket, with no inference of your own.Capabilities
- Omni-modal: text, images, video, audio in one space
- Backwards-compatible with jina-v5-text indexes
- ~239M parameters for edge/high-throughput deployment
- Matryoshka dimensions for flexible storage
- Apache 2.0 license
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| Cross-modal retrieval | Recall@10 | Competitive with 677M variant | Jina AI, May 2026 |
Performance
Common Pipeline Companions
Explore on Mixpeek
Compare alternatives in this category
Hand-picked tools & platforms compared
Deep-dive technical guide
See how Mixpeek runs models as extractors
Store & search embeddings at scale
Usage-based pricing for pipelines
Compare models, APIs & infrastructure
Specification
Research Paper
Jina Embeddings v5 Omni: Multimodal Embeddings for Text, Image, Audio, and Video
arxiv.orgBuild a pipeline with jina-embeddings-v5-omni-nano
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free