embeddinggemma-2
by google
Google's open 740M embedding model that puts text, code, images, video and audio in one 768-d space, small enough for a phone
google/embeddinggemma-2Deploy embeddinggemma-2
Single-tenantMixpeek has no managed extractor for this model. On a single-tenant deployment you upload the weights and a custom plugin serves them next to the rest of your pipeline.
Overview
EmbeddingGemma 2 is Google DeepMind's open multimodal embedding model, released in September 2026 under Apache-2.0. It maps text (including code), images, video and audio, alone or mixed in one input, into a single 768-dimension vector space, so a text query, a photo or a voice memo can all search the same index.
It is built to run on phones and laptops. The text model is 270M parameters, and the vision (170M) and audio (300M) encoders load only if you need them, for 740M in total. It keeps EmbeddingGemma's multilingual text quality (61.36 on MTEB multilingual) and raises code retrieval from 68.76 to 78.68.
Architecture
A Gemma 4-based encoder: 24 layers, model dimension 512, sliding-window attention (1024 tokens) with one global layer per five local, mean pooling and a 512 to 768 projection. Separate vision and audio encoders feed the same backbone, and placeholder tokens (<|image|>, <|video|>, <|audio|>) mark where each item sits in an interleaved input. Trained with Matryoshka Representation Learning and with short task prefixes for queries and documents.
Mixpeek SDK Integration
# Embed a clip and a caption on your side, then upsert the vectors into a
# Mixpeek collection with a 768-d vector index. EmbeddingGemma 2 is not a
# managed extractor, so the inference runs where you choose.
import requests, torch
from sentence_transformers import SentenceTransformer
dtype = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float32
model = SentenceTransformer("google/embeddinggemma-2", model_kwargs={"torch_dtype": dtype}) # never float16
clip = model.encode({"text": "<|video|>", "video": "dock-cam-0412.mp4"}, # media take no task prefix
normalize_embeddings=True)
requests.post(
"https://api.mixpeek.com/v1/namespaces/ns_your_namespace/documents/upsert",
headers={"Authorization": "Bearer API_KEY"},
json={"collection_id": "col_your_collection",
"documents": [{"document_id": "dock-cam-0412",
"vectors": {"embeddinggemma2": clip.tolist()},
"payload": {"source_key": "s3://footage/dock-cam-0412.mp4"}}]},
)Capabilities
- Text (including code), images, video and audio in one 768-dimension space, alone or interleaved in a single input
- Load only what you need: 270M for text, 440M with images, 570M with audio, 740M for everything
- Matryoshka truncation to 512, 256 or 128 dimensions (up to 6x less vector storage)
- 8,192-token context shared by all inputs: about 58 video frames at 1 fps or about 5.5 minutes of audio
- Task prefixes for search, question answering, fact checking, code retrieval, classification, clustering and similarity
- Apache-2.0, so commercial use is allowed
Use Cases on Mixpeek
Benchmarks
| Dataset | Metric | Score | Source |
|---|---|---|---|
| MTEB multilingual v2 | Mean(Task) | 61.36 | Model card: google/embeddinggemma-2 (self-reported, 768d; EmbeddingGemma 1: 61.15) |
| MTEB code v1 | Mean(Task), NDCG@10 | 78.68 | Model card (self-reported; EmbeddingGemma 1: 68.76) |
| MIEB lite (images) | Mean(TaskType) | 64.64 | Model card (self-reported) |
| MMEB v2 VisDoc (visual documents) | Mean(Task), NDCG@5 | 67.84 | Model card (self-reported) |
| MMEB v2 Video | Mean(Task), Hit@1 | 50.67 | Model card (self-reported) |
| MSEB retrieval (audio) | Mean(Task), MRR@10 | 69.54 | Model card (self-reported) |
Performance
Google reports about 191 MB of active RAM for the quantized text-only weights and about 567 MB for the full multimodal model on a Pixel 11 Pro. Run it in bfloat16 or float32: in float16 it returns NaN or silently degraded vectors. The card calls 256d close to lossless and says 128d hurts multimodal quality, so test 128d on your own data first. We have not measured it.
Common Pipeline Companions
Frequently Asked Questions
What is EmbeddingGemma 2?
An open embedding model from Google DeepMind that turns text, code, images, video and audio into vectors in one shared space, so any of them can search any other. It has 740M parameters, runs on phones and laptops, and is licensed Apache-2.0.
Can EmbeddingGemma 2 search video and audio?
Yes. Video is read as frames (1 per second by default, 140 tokens each), so one embedding covers about 58 seconds of video; audio costs 25 tokens a second, so one embedding covers about 5.5 minutes. Longer recordings are split into segments and each segment is embedded. The card reports 50.67 Hit@1 on MMEB video and 69.54 MRR@10 on MSEB audio retrieval.
How is EmbeddingGemma 2 different from EmbeddingGemma?
EmbeddingGemma (300M) embeds text only, with a 2K context. EmbeddingGemma 2 adds images, video and audio in the same space, a 4x larger 8K context, a much stronger code score (78.68 against 68.76) and an Apache-2.0 license, while matching v1 on multilingual text.
Can I use EmbeddingGemma 2 commercially?
Yes. It is released under Apache-2.0, unlike the first EmbeddingGemma, which used the Gemma license.
How much memory does EmbeddingGemma 2 need on a phone?
Google reports about 191 MB of active RAM for the quantized text-only weights and about 567 MB for the full multimodal model, measured on a Pixel 11 Pro. Load only the encoders you need to stay near the lower figure.
How do I use EmbeddingGemma 2 with Mixpeek?
Run the model where you want the inference, on a device or your own GPU, and upsert the 768-d vectors into a Mixpeek collection; a retriever then searches them like any other vector index. Mixpeek does not run EmbeddingGemma 2 as a managed extractor today.
Specification
Research Paper
EmbeddingGemma 2: an open, lightweight multimodal embedding model (Google)
arxiv.orgBuild a pipeline with embeddinggemma-2
Add this model to a processing pipeline alongside other extractors. Combine with retrieval stages for end-to-end search.
Run it on your own data, free