NEWVectors or files. Pick a path.Start →
    Search & Discovery
    7 min read
    Updated 2026-10-06

    How Do I Find Where Users Got Stuck Across Dozens of Usability Test Recordings?

    To find where users got stuck across many usability test recordings, index them so each word is joined to the screen it was spoken over, then ask once across every session. Results open at the moment it happened, attributed to a participant, and grouping by participant shows who hit the problem.

    Usability Testing
    UX Research
    Session Recordings
    Video Search
    Research Synthesis

    How do I find where users got stuck across dozens of usability test recordings?



    Index the recordings so the words and the screen are searchable together, then ask the question once across every session. Transcribe each session with word-level timing, cut it into scenes wherever the screen changes, read the text on screen in each scene, and join each word to the screen it was spoken over. A search such as "where did people hesitate on the checkout step" then returns every matching moment, each one attributed to a participant and opening at the second it happened. Group the results by participant to see exactly who hit the problem.

    How do research teams usually find these moments?



    ApproachWhat it findsWhere it stops
    Watching every session and tagging by handEverything, if someone watches closelyHours per session; the tags reflect what the note-taker was looking for
    Searching the transcriptsWhat participants saidMisses silent struggle, and a quote without the screen behind it is hard to interpret
    Searching the transcript and the screen togetherWhat was said and what was on screen at that secondNeeds the recordings indexed and the study, participant and task recorded as fields

    What does the indexing step need to do?



    1. Record the ids. Store the study, participant, task and recording rig with each session. Every question you ask later filters or groups on these, so they are the schema decision that matters. 2. Cut at screen changes. Splitting on a fixed interval lands results in the middle of an action. Cutting where the display changes makes each result a moment with its own screen state. 3. Transcribe with word timing. Word-level timestamps are what let a quote line up with the exact scene. 4. Read the screen. On-screen text (the page title, the error, the button label) is how you search for a step by name. Speech alone never mentions most of it. 5. Join words to screens. Attach each timed word to the scene it falls in. That join is what lets you check "they said it was confusing" against what they were looking at when they said it.

    How do I ask questions across every session?



    Two kinds of question cover most synthesis work. Find the moment returns the best-matching scenes, for example "the moment someone says they are confused", each opening at its timestamp. Roll up a task groups results by participant for one task, for example "what did everyone say about the new navigation", so every participant appears in the answer with their own moments. Grouping on the participant field matters: without it, the five most similar quotes can all come from the two most talkative people, and the quiet participants disappear from the finding.

    Keep every claim linked to its moment. A generated summary of twelve sessions reads well and cannot be checked. Mixpeek's design notes put it this way: systems that act on unstructured content need verifiable facts derived from the real recordings. A finding that opens at the second it came from is one a stakeholder can verify in a few seconds.

    What results should I expect?



    Measured on our own reference corpus of six moderated sessions:

  1. 7,257 word-timed words joined to 319 screen states.
  2. A median scene of 5 seconds when sessions are cut at screen transitions.
  3. 112 moments where a participant made a claim, 22 setting changes read off the display, and 9 task boundaries
  4. taken from the moderator narrating the protocol.
  5. One measurement worked on only 2 of 6 sessions: timing how long a participant's hands were off a control,
  6. because on the other rigs the camera was framed on the display with the control out of shot. Check your camera framing before you plan a metric that depends on it.

    How do I do this with Mixpeek?



    Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. The UX session analysis template sets up the pipeline above: sessions with study, participant, rig and task fields, scenes cut at screen transitions with transcription, on-screen text and a multimodal embedding per scene, a find-the-moment retriever, and a task roll-up that groups by participant. The recordings stay in your storage, which matters when session footage cannot leave your environment.

    It suits research teams running moderated studies with more sessions than anyone can rewatch. If you run a handful of sessions a quarter, watching them is still the better method. Product analytics on live traffic is a job for a session replay tool. Video is billed per minute on the rate card. Related: how to find one specific moment in hours of footage and why video search misses on-screen text.

    Frequently Asked Questions



    How do I find where users got stuck across dozens of usability test recordings?



    Index the recordings with word-timed transcripts, scenes cut at screen changes and the text on screen, joined together, then search all sessions at once with a question such as "where did people hesitate on checkout". Each result opens at the moment it happened and names the participant.

    Can I search usability recordings for silent struggle?



    Partly. Because each scene carries the screen state, you can search for what was on screen, such as an error message or a step that repeats, as well as for words. Long pauses and repeated attempts are easier to find when scenes are cut at screen changes.

    How do I make sure no participant is left out of a finding?



    Group results by the participant field for each task. A plain search returns the most similar quotes, which often come from the few participants who talked most.

    Do I need to upload session recordings to a third-party tool?



    No. A pipeline that reads from your own storage can index the recordings where they are, which helps when consent terms or policy keep footage inside your environment.
    Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs