How do I search photos, video and audio on the device, without uploading them?
Run an embedding model on the device, store its vectors in a local index, and search that index. A multimodal embedding model turns each photo, each short stretch of video and each audio clip into a vector in one shared space, and a text query or an example becomes a vector in the same space, so "the clip where the dog jumps into the lake" finds the right video without anything leaving the phone or laptop. Google's EmbeddingGemma 2 (September 2026, Apache-2.0) is an open model built for exactly this: text, images, video and audio in one 768-dimension space, in about 191 MB of RAM for text and 567 MB with every encoder loaded, on a Pixel 11 Pro with quantization.
What does the pipeline look like?
1. Pick what to embed. A photo is one input. Video and audio are cut into segments, because one embedding has a fixed budget (see below). 2. Embed each item locally. Load only the encoders you need: text alone, text and images, or everything. 3. Store the vectors in a local index. A small vector store inside the app (an embedded database or a flat file of vectors) is enough for tens of thousands of items. Keep the file path and the segment's start time next to each vector. 4. Search. Embed the query, whether text, a photo or a voice memo, and return the closest vectors with their file and timestamp. 5. Re-embed only what changes. New photos and recordings are embedded as they arrive; nothing else is redone.
How long can a clip be?
Everything in one input shares a budget. In EmbeddingGemma 2 that budget is 8,192 tokens:
| Input | Cost | Most that fits in one embedding |
| Text | 1 token per subword | 8,192 tokens |
| Image | 280 tokens (default) | about 29 images |
| Video | 140 tokens per frame, 1 frame per second by default | about 58 seconds |
| Audio | 25 tokens per second, mono 16 kHz | about 5.5 minutes |
How much storage do the vectors need?
A 768-dimension vector in 32-bit floats is 3,072 bytes. Models trained with Matryoshka representation learning, EmbeddingGemma 2 among them, let you keep only the first 512, 256 or 128 numbers and re-normalize. At 256 dimensions a vector is 1,024 bytes, so 100,000 segments take about 100 MB, and Google's card reports quality close to the full vector at that size. Below 256, multimodal quality drops sharply on the card's benchmarks, so test it on your own library first. Queries and the library must use the same size.
What goes wrong most often?
When does it stop fitting on the device?
On-device search suits one person's library on one device. It gets harder when the library is shared by a team, lives in cloud storage already, or grows past what a phone should store and re-embed. That is the point where the same embeddings move to a server-side index.
Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. Video is indexed by scene with timestamps, starting at $0.05 a minute on the rate card. If you already produce EmbeddingGemma 2 vectors, MVS stores and searches your own vectors on object storage. Mixpeek does not run EmbeddingGemma 2 as a managed extractor today. Related: how to find one specific moment in hours of footage and the EmbeddingGemma 2 model page.
Frequently Asked Questions
How do I search photos, video and audio on the device, without uploading them?
Embed each photo, video segment and audio clip with a multimodal embedding model running locally, store the vectors in a local index, and search it by embedding the query the same way. EmbeddingGemma 2 is an open model built for this, at about 567 MB of RAM with all encoders loaded.
Can I search video with a voice memo?
Yes, with a model that puts audio and video in one space. The voice memo is embedded as audio and compared with the video segments' embeddings, so the closest segments come back with their timestamps.
How long a video can one embedding cover?
With EmbeddingGemma 2's defaults, about 58 seconds: one frame per second at 140 tokens a frame within an 8,192-token budget. Longer videos are split into segments.
How much does on-device multimodal search cost?
The model is free under Apache-2.0 and runs on the device, so there is no per-item fee. The costs are device memory (about 191 MB for text only, 567 MB for everything) and storage, about 1 KB per item at 256 dimensions.