NEWVectors or files. Pick a path.Start →
    Search & Discovery
    7 min read
    Updated 2026-10-06

    How Do I Search Photos, Video and Audio on the Device, Without Uploading Them?

    To search photos, video and audio on a phone or laptop without uploading them, embed each item with a multimodal model running on the device, keep the vectors in a local index, and search it with text, a photo or a voice memo. EmbeddingGemma 2 does this in about 567 MB of RAM.

    On-Device AI
    Multimodal Embeddings
    EmbeddingGemma
    Video Search
    Privacy

    How do I search photos, video and audio on the device, without uploading them?



    Run an embedding model on the device, store its vectors in a local index, and search that index. A multimodal embedding model turns each photo, each short stretch of video and each audio clip into a vector in one shared space, and a text query or an example becomes a vector in the same space, so "the clip where the dog jumps into the lake" finds the right video without anything leaving the phone or laptop. Google's EmbeddingGemma 2 (September 2026, Apache-2.0) is an open model built for exactly this: text, images, video and audio in one 768-dimension space, in about 191 MB of RAM for text and 567 MB with every encoder loaded, on a Pixel 11 Pro with quantization.

    What does the pipeline look like?



    1. Pick what to embed. A photo is one input. Video and audio are cut into segments, because one embedding has a fixed budget (see below). 2. Embed each item locally. Load only the encoders you need: text alone, text and images, or everything. 3. Store the vectors in a local index. A small vector store inside the app (an embedded database or a flat file of vectors) is enough for tens of thousands of items. Keep the file path and the segment's start time next to each vector. 4. Search. Embed the query, whether text, a photo or a voice memo, and return the closest vectors with their file and timestamp. 5. Re-embed only what changes. New photos and recordings are embedded as they arrive; nothing else is redone.

    How long can a clip be?



    Everything in one input shares a budget. In EmbeddingGemma 2 that budget is 8,192 tokens:

    InputCostMost that fits in one embedding
    Text1 token per subword8,192 tokens
    Image280 tokens (default)about 29 images
    Video140 tokens per frame, 1 frame per second by defaultabout 58 seconds
    Audio25 tokens per second, mono 16 kHzabout 5.5 minutes
    So a 20-minute video becomes about 25 segments of 45 to 50 seconds each, and each search result points to the segment with its start time. Shorter segments find moments more precisely and cost more embeddings.

    How much storage do the vectors need?



    A 768-dimension vector in 32-bit floats is 3,072 bytes. Models trained with Matryoshka representation learning, EmbeddingGemma 2 among them, let you keep only the first 512, 256 or 128 numbers and re-normalize. At 256 dimensions a vector is 1,024 bytes, so 100,000 segments take about 100 MB, and Google's card reports quality close to the full vector at that size. Below 256, multimodal quality drops sharply on the card's benchmarks, so test it on your own library first. Queries and the library must use the same size.

    What goes wrong most often?



  1. Running in float16. EmbeddingGemma 2 overflows float16 and returns NaN or quietly worse vectors with no
  2. error. Use bfloat16 where the hardware supports it and float32 elsewhere.
  3. Skipping the task prefix. Text queries embed better with the search prefix, and documents with
  4. "title: ... | text: ...". Images, video and audio take no prefix.
  5. One embedding for a whole long video. It will not fit the budget, and even when it does, a single vector
  6. cannot tell you where in the video the match is. Segment it.
  7. Expecting text and images to score alike. Text-to-text matches score higher than text-to-image matches in
  8. most shared spaces, so a mixed result list ranks text first. The modality gap explains why and how to rebalance.

    When does it stop fitting on the device?



    On-device search suits one person's library on one device. It gets harder when the library is shared by a team, lives in cloud storage already, or grows past what a phone should store and re-embed. That is the point where the same embeddings move to a server-side index.

    Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. Video is indexed by scene with timestamps, starting at $0.05 a minute on the rate card. If you already produce EmbeddingGemma 2 vectors, MVS stores and searches your own vectors on object storage. Mixpeek does not run EmbeddingGemma 2 as a managed extractor today. Related: how to find one specific moment in hours of footage and the EmbeddingGemma 2 model page.

    Frequently Asked Questions



    How do I search photos, video and audio on the device, without uploading them?



    Embed each photo, video segment and audio clip with a multimodal embedding model running locally, store the vectors in a local index, and search it by embedding the query the same way. EmbeddingGemma 2 is an open model built for this, at about 567 MB of RAM with all encoders loaded.

    Can I search video with a voice memo?



    Yes, with a model that puts audio and video in one space. The voice memo is embedded as audio and compared with the video segments' embeddings, so the closest segments come back with their timestamps.

    How long a video can one embedding cover?



    With EmbeddingGemma 2's defaults, about 58 seconds: one frame per second at 140 tokens a frame within an 8,192-token budget. Longer videos are split into segments.

    How much does on-device multimodal search cost?



    The model is free under Apache-2.0 and runs on the device, so there is no per-item fee. The costs are device memory (about 191 MB for text only, 567 MB for everything) and storage, about 1 KB per item at 256 dimensions.
    Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs