NEWVectors or files. Pick a path.Start →
    Back to DiagramsIR Foundations

    You search inside the file, not for it

    You search inside the file, not for it

    Object decomposition: you search inside the file, not for it. A one-hour recording.mp4 is split into scenes (visual embeddings), speech (transcript embeddings on silence and speaker turns), faces (who appeared and when), and on-screen text (OCR of slides and captions). A whole file being findable is not the same as usable — the thing you need is a moment inside it, like 23:41 where the renewal number is spoken while the pricing slide is on screen. The same move works on video, PDFs, and audio: one file becomes dozens of small, independently searchable documents.
    You search inside the file, not for it

    The biggest unlock in multimodal search isn't a model. It's a decision: stop treating the file as the unit of search.

    A one-hour recording stored as one record can be found, but not used. The thing you actually need is at 23:41, where someone says the renewal number while the pricing slide is on screen. So you decompose: split on scene changes, silence, speaker turns. Each segment gets its own embeddings for what was shown, what was said, who appeared, and what text was on screen.

    One file becomes dozens of small searchable documents. The same move works on PDFs (pages, then paragraphs) and audio (transcript segments).

    This is the first thing Mixpeek does to any file you ingest, and every downstream capability depends on it.

    Nobody ever needed the video. They needed the moment at 23:41.

    Run this on your own data

    Mixpeek turns video, images, audio, and documents in your object storage into searchable, timestamped results through one API.

    Search your own data, free