Why does AI miss things when I ask it about a long video?
The model only sees some of the frames. A model that accepts video samples frames from it, and a model that accepts images only sees the frames you send, so the moment you are asking about can fall between them. Gemini samples one frame per second by default, and its documentation notes that fast action can lose detail at that rate. At one frame per second, an hour of video is 3,600 frames, more than Claude's API takes in one request (up to 600 images) or OpenAI's (up to 1,500). The fix that holds up is to find the minutes your question is about first, then send only those to the model at a higher frame rate.
How much of a long video does the model actually see?
What each API documents for video input, and what a one-hour video becomes under those limits. The last column is arithmetic on the documented numbers.
| API | How video goes in | Documented limit | A one-hour video |
| Gemini API | Video file, sampled at 1 frame per second by default; a custom fps can be set | 258 tokens per frame (66 at low media resolution) plus 32 audio tokens per second; a 1M-token context holds about 1 hour at 258 tokens per frame or 3 hours at low | 3,600 frames, about 1.04M tokens with audio at 258 tokens per frame |
| Claude API | Images only, no video input | 600 images per request (100 on models with a 200k context, 20 per message in claude.ai); a frame costs ceil(width/28) x ceil(height/28) tokens | At most one frame every 6 seconds; a 1280x720 frame is 1,196 tokens, so 600 frames is about 718,000 |
| OpenAI API | Images only in the vision guide, which lists no video input | 1,500 images per request, 512 MB per request | At most one frame every 2.4 seconds |
Why do short events get missed?
At one frame per second, anything shorter than a second can land between two samples: a logo on screen for half a second, a jersey number in a fast pan, a one-shot cutaway. When you pick the frames yourself to stay under an image cap, the gap grows. Spread 600 frames across an hour and the model sees one every six seconds.
Sound is a separate channel. Gemini reads the audio track with the video. An API that takes images hears nothing, so anything that was only said is gone unless you add a transcript as text.
Why do answers get vaguer as the video gets longer?
A longer input puts more frames in front of the model for the same question, and most of them are unrelated to it. Research on long text inputs found that models use information at the start and end of the input more reliably than information in the middle (Liu et al., "Lost in the Middle", 2023). A detail from minute 34 of a 60-minute video sits in the middle of the input.
Cost moves the same way. Every frame is billed as input tokens on every question you ask, so asking ten questions about an hour of video pays for that hour ten times.
How do I get reliable answers about a long video?
Search first, then ask.
1. Split the video into segments with timestamps. Cut at scene changes for edited footage and at fixed windows for lectures, meetings or camera feeds. The video scene segmentation guide covers the trade-offs. 2. Index each segment. Store an embedding of its frames, its transcript, and any text shown on screen. On-screen text needs its own step; why video search misses on-screen text explains why. 3. Search for the part your question is about. Run the question against the index and keep the top few segments, with their start and end times. 4. Send only those segments to the model, sampled densely. A two-minute clip at several frames per second shows the model far more of that moment than the whole hour at one frame per second. 5. Ask with the timestamps attached, so the answer can say where in the video it found each claim and you can check it.
For a question about the whole video, such as a summary, summarize each segment first and then summarize those summaries. Each call then sees a short span in full.
When is it fine to send the whole video?
When the video is short, when the question is about the overall gist, or when you need the model to follow events across the entire timeline in one pass. Gemini can take about an hour in one request at 258 tokens per frame, and about three at low media resolution. Once you ask about specific moments, or ask many questions about the same footage, searching first is more accurate and cheaper.
How do I do this with Mixpeek?
Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. The video scene search recipe does steps 1 to 3: it splits each video at scene changes, embeds each scene's frames and transcript, and returns each match as the video's ID with the scene's start and end time. Cut those spans out and send them to the model with your question.
import subprocess
from mixpeek import Mixpeek
client = Mixpeek(api_key="YOUR_API_KEY", namespace="video-scenes")
# The retriever from the video scene search recipe
results = client.retrievers.execute(
"ret_your_scene_retriever",
inputs={"query": "the speaker walks through the revenue chart"},
)
for i, doc in enumerate(results["documents"][:3]):
start, end = doc["start_time"], doc["end_time"]
# Re-encode so the clip starts exactly at the scene boundary
subprocess.run(
["ffmpeg", "-y", "-ss", str(start), "-to", str(end), "-i", "talk.mp4", f"clip_{i}.mp4"],
check=True,
)
print(f"clip_{i}.mp4 covers {start}s to {end}s")
# Send clip_0.mp4 to clip_2.mp4 to your vision model with the question and these timestampsFrequently Asked Questions
Why does AI miss things when I ask it about a long video?
It only sees sampled frames. Gemini reads one frame per second by default, and image-only APIs see just the frames you send, capped at 600 per request on Claude and 1,500 on OpenAI. Anything between those frames, or anything only spoken when no transcript is sent, is invisible to the model.
Can ChatGPT or Claude watch a whole video?
Their APIs take images, so you send frames from the video. Claude's API takes up to 600 images per request and OpenAI's up to 1,500, which across an hour of video is one frame every 6 or 2.4 seconds at best. Gemini's API takes a video file directly and samples it at one frame per second.
How many frames per second does Gemini read from a video?
One frame per second by default, at 258 tokens per frame (66 at low media resolution) plus 32 tokens per second of audio. You can set a custom frame rate, and the documentation notes that fast action can lose detail at the default.
What is the best way to ask questions about a long video?
Find the part of the video your question is about first, then send only that part to the model at a higher frame rate, with its timestamps. Splitting the video into segments, indexing each one, and searching that index does the finding.