How do I find one specific moment in hours of footage?
Start from what you remember about the moment, because each kind of clue has a faster route than scrubbing. If you know roughly when it happened, filter by timecode and file dates. If you remember something that was said, search a timestamped transcript. If you remember what was on screen, use a visual search that lets you describe the scene in words. If it was a sound, such as applause, a siren or a goal horn, search for the sound. For one long recording you can do most of this in your editing software; for a library of footage you need an index built once, so every later search takes seconds.
What do you remember about the moment?
| What you remember | Fastest way to find it | Works on |
| Roughly when it happened | Sort and filter by timecode, file creation time or camera metadata, then scrub only that window | Any footage with intact metadata |
| Something that was said | Transcribe with timestamps, then search the transcript and jump to the time | Speech in any language the transcriber supports |
| Words that appeared on screen | OCR on sampled frames, then search the text | Titles, signs, slides, scoreboards, captions |
| What it looked like | Describe it in words ("a man in a red jacket opens a door at night") and search visual embeddings of each scene | Anything visible |
| A sound | Detect audio events (laughter, applause, sirens, music) and search by event | Non-speech audio |
| A similar shot you already have | Search by example: use a frame or clip as the query | Re-used shots, re-edits, B-roll |
How do I find a moment in one long recording?
For "somewhere in the last three months of shoots", you need an index, covered next.
How do I find a moment across a whole library of footage?
Index the library once, then search it. Indexing means:
1. Split every video into scenes or short segments, each with a start and end time, so a search result lands on a moment inside the file. 2. Describe each segment several ways: a transcript of what was said, text read off the screen, detected audio events, and a visual embedding that captures what the scene looks like. 3. Store those with the segment, along with the file path and timestamps, in a search index. 4. Search in plain language and get back ranked segments that open at the right second.
The cost is a one-time processing pass over the footage; after that a search takes well under a second. What each part costs is broken down in how much it costs to make a media library searchable.
Why does the search still return the wrong moment?
Usually because segments are too long, so one result spans several scenes, or because the search only reads the transcript and your clue was visual. Why does my video search return the wrong scene covers both, and why video search misses on-screen text covers words that are shown but never spoken.
How do I do this with Mixpeek?
Mixpeek indexes footage where it already lives, in S3, GCS or any bucket. The multimodal extractor splits each video by time, scene or silence and stores every segment with its transcript, on-screen text, a description and a visual embedding. A retriever then searches all of them at once and returns segments with their start and end times, so a result opens at the moment. The Footage Intelligence template sets up that pipeline in one click, and the video understanding tutorial walks through it step by step.
Related: how to search meeting recordings for what someone said, the best video search tools and how AI finds specific moments in video, the research behind this.
Frequently Asked Questions
How do I find one specific moment in hours of footage?
Start from what you remember. Filter by timecode if you know when it happened, search a timestamped transcript for something that was said, use OCR for words on screen, describe the scene in words for something you saw, or search audio events for a sound. For a whole library, index the footage once into timestamped segments so every search returns the moment directly.
Can I search video by describing what happens in it?
Yes. Visual embedding models place a text description and a video frame or clip in the same space, so "a dog jumping into a pool" matches scenes that show it even when nobody says those words. The footage has to be indexed first, segment by segment.
What if the moment has no dialogue?
Use the other clues: what was on screen (visual search or OCR) and what it sounded like (audio event detection). Transcript search only finds moments where someone spoke.
How long does it take to index hours of footage?
It runs once per file and scales with the amount of video and the number of things extracted from it. After indexing, searches return in well under a second regardless of library size.