NEWVectors or files. Pick a path.Start →
    Search & Discovery
    7 min read
    Updated 2026-10-10

    Why Does AI Miss Things When I Ask It About a Long Video?

    AI misses things in a long video because it only sees sampled frames. Gemini reads one frame per second by default, and Claude's and OpenAI's APIs take images only, capped at 600 and 1,500 per request, while an hour of video at one frame per second is 3,600 frames. The fix is to find the minutes your question is about first and send only those, at a higher frame rate.

    Video Understanding
    Vision Language Models
    Long Video
    Frame Sampling
    Video Search

    Why does AI miss things when I ask it about a long video?



    The model only sees some of the frames. A model that accepts video samples frames from it, and a model that accepts images only sees the frames you send, so the moment you are asking about can fall between them. Gemini samples one frame per second by default, and its documentation notes that fast action can lose detail at that rate. At one frame per second, an hour of video is 3,600 frames, more than Claude's API takes in one request (up to 600 images) or OpenAI's (up to 1,500). The fix that holds up is to find the minutes your question is about first, then send only those to the model at a higher frame rate.

    How much of a long video does the model actually see?



    What each API documents for video input, and what a one-hour video becomes under those limits. The last column is arithmetic on the documented numbers.

    APIHow video goes inDocumented limitA one-hour video
    Gemini APIVideo file, sampled at 1 frame per second by default; a custom fps can be set258 tokens per frame (66 at low media resolution) plus 32 audio tokens per second; a 1M-token context holds about 1 hour at 258 tokens per frame or 3 hours at low3,600 frames, about 1.04M tokens with audio at 258 tokens per frame
    Claude APIImages only, no video input600 images per request (100 on models with a 200k context, 20 per message in claude.ai); a frame costs ceil(width/28) x ceil(height/28) tokensAt most one frame every 6 seconds; a 1280x720 frame is 1,196 tokens, so 600 frames is about 718,000
    OpenAI APIImages only in the vision guide, which lists no video input1,500 images per request, 512 MB per requestAt most one frame every 2.4 seconds
    Sources: Gemini video understanding, Claude vision, OpenAI images and vision. Checked 2026-10-10.

    Why do short events get missed?



    At one frame per second, anything shorter than a second can land between two samples: a logo on screen for half a second, a jersey number in a fast pan, a one-shot cutaway. When you pick the frames yourself to stay under an image cap, the gap grows. Spread 600 frames across an hour and the model sees one every six seconds.

    Sound is a separate channel. Gemini reads the audio track with the video. An API that takes images hears nothing, so anything that was only said is gone unless you add a transcript as text.

    Why do answers get vaguer as the video gets longer?



    A longer input puts more frames in front of the model for the same question, and most of them are unrelated to it. Research on long text inputs found that models use information at the start and end of the input more reliably than information in the middle (Liu et al., "Lost in the Middle", 2023). A detail from minute 34 of a 60-minute video sits in the middle of the input.

    Cost moves the same way. Every frame is billed as input tokens on every question you ask, so asking ten questions about an hour of video pays for that hour ten times.

    How do I get reliable answers about a long video?



    Search first, then ask.

    1. Split the video into segments with timestamps. Cut at scene changes for edited footage and at fixed windows for lectures, meetings or camera feeds. The video scene segmentation guide covers the trade-offs. 2. Index each segment. Store an embedding of its frames, its transcript, and any text shown on screen. On-screen text needs its own step; why video search misses on-screen text explains why. 3. Search for the part your question is about. Run the question against the index and keep the top few segments, with their start and end times. 4. Send only those segments to the model, sampled densely. A two-minute clip at several frames per second shows the model far more of that moment than the whole hour at one frame per second. 5. Ask with the timestamps attached, so the answer can say where in the video it found each claim and you can check it.

    For a question about the whole video, such as a summary, summarize each segment first and then summarize those summaries. Each call then sees a short span in full.

    When is it fine to send the whole video?



    When the video is short, when the question is about the overall gist, or when you need the model to follow events across the entire timeline in one pass. Gemini can take about an hour in one request at 258 tokens per frame, and about three at low media resolution. Once you ask about specific moments, or ask many questions about the same footage, searching first is more accurate and cheaper.

    How do I do this with Mixpeek?



    Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. The video scene search recipe does steps 1 to 3: it splits each video at scene changes, embeds each scene's frames and transcript, and returns each match as the video's ID with the scene's start and end time. Cut those spans out and send them to the model with your question.
    import subprocess
    from mixpeek import Mixpeek
    
    client = Mixpeek(api_key="YOUR_API_KEY", namespace="video-scenes")
    
    # The retriever from the video scene search recipe
    results = client.retrievers.execute(
        "ret_your_scene_retriever",
        inputs={"query": "the speaker walks through the revenue chart"},
    )
    
    for i, doc in enumerate(results["documents"][:3]):
        start, end = doc["start_time"], doc["end_time"]
        # Re-encode so the clip starts exactly at the scene boundary
        subprocess.run(
            ["ffmpeg", "-y", "-ss", str(start), "-to", str(end), "-i", "talk.mp4", f"clip_{i}.mp4"],
            check=True,
        )
        print(f"clip_{i}.mp4 covers {start}s to {end}s")
    
    # Send clip_0.mp4 to clip_2.mp4 to your vision model with the question and these timestamps
    Mixpeek fits teams with a video library in object storage who ask many questions of the same footage. It is more than you need for asking one question about one short clip; send that clip straight to the model. Prices are on the pricing page. For choosing the model that reads the clips, see the best vision language models and best AI video analysis tools. For the research side of long video, see long-context video understanding.

    Frequently Asked Questions



    Why does AI miss things when I ask it about a long video?



    It only sees sampled frames. Gemini reads one frame per second by default, and image-only APIs see just the frames you send, capped at 600 per request on Claude and 1,500 on OpenAI. Anything between those frames, or anything only spoken when no transcript is sent, is invisible to the model.

    Can ChatGPT or Claude watch a whole video?



    Their APIs take images, so you send frames from the video. Claude's API takes up to 600 images per request and OpenAI's up to 1,500, which across an hour of video is one frame every 6 or 2.4 seconds at best. Gemini's API takes a video file directly and samples it at one frame per second.

    How many frames per second does Gemini read from a video?



    One frame per second by default, at 258 tokens per frame (66 at low media resolution) plus 32 tokens per second of audio. You can set a custom frame rate, and the documentation notes that fast action can lose detail at the default.

    What is the best way to ask questions about a long video?



    Find the part of the video your question is about first, then send only that part to the model at a higher frame rate, with its timestamps. Splitting the video into segments, indexing each one, and searching that index does the finding.
    Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs