How do I find where users got stuck across dozens of usability test recordings?
Index the recordings so the words and the screen are searchable together, then ask the question once across every session. Transcribe each session with word-level timing, cut it into scenes wherever the screen changes, read the text on screen in each scene, and join each word to the screen it was spoken over. A search such as "where did people hesitate on the checkout step" then returns every matching moment, each one attributed to a participant and opening at the second it happened. Group the results by participant to see exactly who hit the problem.
How do research teams usually find these moments?
| Approach | What it finds | Where it stops |
| Watching every session and tagging by hand | Everything, if someone watches closely | Hours per session; the tags reflect what the note-taker was looking for |
| Searching the transcripts | What participants said | Misses silent struggle, and a quote without the screen behind it is hard to interpret |
| Searching the transcript and the screen together | What was said and what was on screen at that second | Needs the recordings indexed and the study, participant and task recorded as fields |
What does the indexing step need to do?
1. Record the ids. Store the study, participant, task and recording rig with each session. Every question you ask later filters or groups on these, so they are the schema decision that matters. 2. Cut at screen changes. Splitting on a fixed interval lands results in the middle of an action. Cutting where the display changes makes each result a moment with its own screen state. 3. Transcribe with word timing. Word-level timestamps are what let a quote line up with the exact scene. 4. Read the screen. On-screen text (the page title, the error, the button label) is how you search for a step by name. Speech alone never mentions most of it. 5. Join words to screens. Attach each timed word to the scene it falls in. That join is what lets you check "they said it was confusing" against what they were looking at when they said it.
How do I ask questions across every session?
Two kinds of question cover most synthesis work. Find the moment returns the best-matching scenes, for example "the moment someone says they are confused", each opening at its timestamp. Roll up a task groups results by participant for one task, for example "what did everyone say about the new navigation", so every participant appears in the answer with their own moments. Grouping on the participant field matters: without it, the five most similar quotes can all come from the two most talkative people, and the quiet participants disappear from the finding.
Keep every claim linked to its moment. A generated summary of twelve sessions reads well and cannot be checked. Mixpeek's design notes put it this way: systems that act on unstructured content need verifiable facts derived from the real recordings. A finding that opens at the second it came from is one a stakeholder can verify in a few seconds.
What results should I expect?
Measured on our own reference corpus of six moderated sessions:
How do I do this with Mixpeek?
Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. The UX session analysis template sets up the pipeline above: sessions with study, participant, rig and task fields, scenes cut at screen transitions with transcription, on-screen text and a multimodal embedding per scene, a find-the-moment retriever, and a task roll-up that groups by participant. The recordings stay in your storage, which matters when session footage cannot leave your environment.
It suits research teams running moderated studies with more sessions than anyone can rewatch. If you run a handful of sessions a quarter, watching them is still the better method. Product analytics on live traffic is a job for a session replay tool. Video is billed per minute on the rate card. Related: how to find one specific moment in hours of footage and why video search misses on-screen text.
Frequently Asked Questions
How do I find where users got stuck across dozens of usability test recordings?
Index the recordings with word-timed transcripts, scenes cut at screen changes and the text on screen, joined together, then search all sessions at once with a question such as "where did people hesitate on checkout". Each result opens at the moment it happened and names the participant.
Can I search usability recordings for silent struggle?
Partly. Because each scene carries the screen state, you can search for what was on screen, such as an error message or a step that repeats, as well as for words. Long pauses and repeated attempts are easier to find when scenes are cut at screen changes.
How do I make sure no participant is left out of a finding?
Group results by the participant field for each task. A plain search returns the most similar quotes, which often come from the few participants who talked most.
Do I need to upload session recordings to a third-party tool?
No. A pipeline that reads from your own storage can index the recordings where they are, which helps when consent terms or policy keep footage inside your environment.