NEWVectors or files. Pick a path.Start →
    Media
    Template v1.0 · updated 2026-09-13

    UX session analysis

    Moderated sessions become a corpus you can question. The transcript is joined to what was on screen at that second, so one query reaches every session at once and returns the moment rather than the recording.

    Ask one question across every session and get every participant's answer, attributed. The query that defines this template is the roll-up: one task, every participant, with each result opening at the second it happened.

    UX research teams running moderated usability sessions

    What deploys today5 ready1 need your input

    No marketplace starter set ships with this template: it applies empty and reads Three to five sessions from ONE study, so connect your own data before querying it.

    7,257 / 319
    word-timed words joined to screen states
    Measured on the reference corpus of six sessions. The join is the product: a word without the screen it was spoken over is a transcript, and a screen state without the word is a screenshot.
    5s
    median scene, cut at screen transitions
    Segmenting on the display changing rather than on a fixed interval is what makes a result land on a behaviour instead of in the middle of one.
    112 / 22 / 9
    claim moments, setting changes, task boundaries
    From the same six sessions: 112 moments where a participant made a claim, 22 setting changes with durations read off the display rather than from a knob sensor, and 9 task boundaries taken from the moderator narrating the protocol.
    2 of 6
    sessions where hand-off-wheel timing works
    61.5% of frames on one rig and 15.9% on another, agreeing with 7 of 9 hand-labelled frames. It fails on the rest because the camera is framed on the display with the wheel out of shot. This is a rig property, not a model property.

    What deploys today

    session-scenesCollectionDeploys ready
    session-transcriptsCollectionDeploys ready
    find-the-momentRetrieverDeploys ready
    task-roll-upRetrieverDeploys ready
    group_byRetriever stageDeploys ready
    clinic-recordingsData sourceNeeds your input
    What each part needs

    The states are read from the manifest, so a part it ships commented out never shows as ready.

    session-scenes
    Deploys and processes documents as they arrive.
    session-transcripts
    Deploys and processes documents as they arrive.
    find-the-moment
    Deploys and answers queries once its collections hold documents.
    task-roll-up
    Deploys and answers queries once its collections hold documents.
    group_by
    The one-click manifest includes this stage.
    clinic-recordings
    Create the s3 credentials as manifest secrets, then uncomment the sync. It ships commented out because apply tests a connection as it creates it.

    What it looks like

    Pick a frame or a search and see what fires.

    Simulated walkthrough · illustrative frames, scripted decisions
    Two people at a wall-mounted touchscreen, one reaching toward the display while the other watchescleared
    session-scenes -> session-transcriptsStill: Pexels / Darlene Aldersonframe 1 of 1
    Detections on this frame
    Decision path
    1. session-scenes -> session-transcripts
    2. participant
    3. screen state at that second
    4. route: cleared
    cut at the screen transition, transcript joined to it
    Running tally
    1
    cleared
    0
    review
    0
    excluded
    What a person's call does here
    None. No saves, no flags, no observer marks are recorded by the template as shipped.
    Try a search
    Pick a search to see which retriever answers it and whether it works on a one-click deploy.
    Where the decisions fire
    Illustrative frames; the boxes are authored to show the decision path
    Two people at a wall-mounted touchscreen, one reaching toward the display while the other watchesparticipantscreen state at that secondcleared
    session-scenes -> session-transcripts · The spoken word and the display are indexed together, so one query reaches the second a participant hit the problem and shows what was on screen when they did.Still: Pexels / Darlene Alderson
    What goes in
    Connect your data to get started.
    sessions
    One object per recording, with study_id, participant_id, rig_id and task_id as fields. Those four are what every roll-up groups and filters on, so they are the schema decision that matters.
    What comes out
    Outputs from this template.
    session-scenes
    Deploys ready
    Segment documents with a 1408-dimension multimodal embedding, timing, transcription, and the on-screen text read off that segment.
    session-transcripts
    Deploys ready
    The spoken side as a 1024-dimension text embedding, carrying the study, participant and task through so a filtered question stays filtered.

    What you would ask it

    Example searches this namespace answers once it is applied. Each one names the retriever that serves it.

    • “where did people hesitate on the checkout step”

      find-the-moment

      the word and the screen state at that second, in one result

    • “what did everyone say about the new navigation”

      task-roll-up

      grouped on participant_id, so no participant is missing from the answer

    • “the moment someone says they are confused”

      find-the-moment

      opens at the second it happened rather than at the recording

    How the namespace is wired

    The diagram shows 1 buckets, 2 collections, 0 clean views and 2 retrievers. The manifest below applies 1 buckets, 2 collections (clean views included) and 2 retrievers today; the other parts are commented out in it, each with the reason. The diagram generates the manifest, so they cannot drift apart.

    SourceBucketCollectionClean viewRetrieverClusterConnectionBucket syncAlertTriggerClick a node to inspect it
    syncsearchsearchpipeline inpipeline in

    Reward signals

    How reviewer decisions move the thresholds

    Thresholds at ingest drift as the corpus changes. The reviewers working the queue are the ones who see where a threshold is wrong first, so this template routes their decisions back into the model that set it.

    Nothing in this template writes signals back. A researcher marking a moment as the right one would be the obvious loop and it is not wired.
    Explicit signals

    None. No saves, no flags, no observer marks are recorded by the template as shipped.

    Implicit signals

    None from the application. Mixpeek records retriever executions server-side, which is telemetry about queries rather than feedback about results.

    Where they land
    mxp_retriever_executions

    System collections in your namespace, on the same vector store as the rest of the template. They are yours to query.

    How the loop closes

    The observers' notes are the ground truth this use case is missing: task start and stop, errors, and confidence per participant. That is an input to evaluation rather than a signal the template collects.

    One file spins up the namespace. Generated from the diagram above. Also served at /templates/ux-session-analysis.namespace.yaml.

    # ux-session-analysis: one manifest spins up the namespace.
    # Platform manifest schema (GET /v1/discovery/schema). Validate with POST /v1/manifest/validate,
    # apply with POST /v1/manifest/apply or the Deploy button. Wiring comes from the flow diagram:
    # edges are bucket -> collection sources, collection -> retriever scope, retriever -> view.
    # Applying a SECOND time, over a namespace this template already created: use
    # POST /v1/manifest/apply?mode=create_missing, which creates what is missing and leaves
    # what exists alone. The default, create_only, fails the WHOLE apply and rolls it back if
    # any resource already exists, so an upgrade looks like a dead end without this. Use
    # mode=upsert to also patch resources that exist but have drifted from this file.
    version: '1.0'
    metadata:
      name: ux-session-analysis
      description: "Namespace template ux-session-analysis. Generated from the flow diagram on mixpeek.com/templates/ux-session-analysis."
    namespaces:
      - name: ux-session-analysis
        description: "Everything below lives in this namespace."
        feature_extractors:
          - name: multimodal_extractor
            version: v1
          - name: text_extractor
            version: v1
        # Daily spend budget, in dollars of what you are charged (1,000 credits a dollar).
        # New batches pause once the namespace has spent this much today; queries are not capped.
        # Raise it later on the namespace: PATCH /v1/namespaces/<id> {"spend_budget": {"daily_usd": N}}.
        budget:
          daily_usd: 0.1
    
    # Data sources. A storage connection carries credentials, so it is created in Studio
    # (or POST /v1/organizations/storage-connections) and synced into the bucket named here.
    #   clinic-recordings: s3, continuous, moderated session recordings from your rig, in your account -> bucket sessions
    buckets:
      - name: sessions
        namespace: ux-session-analysis
        description: "Meant to be fed by clinic-recordings (s3, continuous). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
        schema:
          properties:
            content:
              type: video
              required: true
            study_id:
              type: string
              required: true
            participant_id:
              type: string
              required: true
            task_id:
              type: string
            rig_id:
              type: string
    
    # Connect your own storage. Create the two secrets, uncomment, and apply again;
    # apply live-tests the connection, so it must have real credentials to succeed.
    # Until then the buckets above accept direct uploads.
    # storage_connections:
    #   - name: clinic-recordings-connection
    #     provider: s3
    #     description: "Read-only access to the s3 location holding moderated session recordings from your rig, in your account."
    #     config:
    #       region: us-east-1
    #       credentials:
    #         type: access_key
    #         access_key_id: ${{ secrets.AWS_ACCESS_KEY_ID }}
    #         secret_access_key: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
    
    # bucket_syncs:
    #   - name: clinic-recordings-sync
    #     bucket: sessions
    #     connection: clinic-recordings-connection
    #     source_path: "sessions/"
    #     sync_mode: continuous
    #     polling_interval_seconds: 300
    #     skip_duplicates: true
    #     file_filters:
    #       include_patterns: ["*.mp4", "*.mov"]
    #     schema_mapping:
    #       mappings:
    #         content:
    #           target_type: blob
    #           source: {type: file}
    #           blob_type: video
    
    collections:
      - name: session-scenes
        namespace: ux-session-analysis
        description: "multimodal_extractor@v1 over bucket sessions. Feeds session-transcripts, find-the-moment, task-roll-up."
        source:
          type: bucket
          bucket: sessions
        feature_extractor:
          name: multimodal_extractor
          version: v1
          parameters:
            run_transcription: true
            run_ocr: true
          input_mappings:
            video: content
          field_passthrough:
            - source_path: study_id
            - source_path: participant_id
            - source_path: rig_id
            - source_path: task_id
            - source_path: file_location
            - source_path: segment_id
            - source_path: start_time
            - source_path: end_time
        enabled: true
      - name: session-transcripts
        namespace: ux-session-analysis
        description: "text_extractor@v1 over collection session-scenes. Feeds find-the-moment, task-roll-up."
        source:
          type: collection
          collection: session-scenes
        feature_extractor:
          name: text_extractor
          version: v1
          input_mappings:
            text: transcription
          field_passthrough:
            - source_path: study_id
            - source_path: participant_id
            - source_path: task_id
            - source_path: segment_id
            - source_path: video_segment_url
            - source_path: start_time
            - source_path: end_time
        enabled: true
    retrievers:
      - name: find-the-moment
        namespace: ux-session-analysis
        description: "Searches session-scenes, session-transcripts across 2 feature indexes with rrf fusion."
        collections:
          - session-scenes
          - session-transcripts
        input_schema:
          query:
            type: text
            required: true
            description: "What to look for; searched across every index below"
        stages:
          - stage_name: search
            stage_id: feature_search
            parameters:
              searches:
                - feature_uri: "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1"
                  query:
                    input_mode: text
                    value: "{{INPUT.query}}"
                  top_k: 25
                - feature_uri: "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"
                  query:
                    input_mode: text
                    value: "{{INPUT.query}}"
                  top_k: 25
              fusion: rrf
              final_top_k: 25
        tags:
          - template:ux-session-analysis
      - name: task-roll-up
        namespace: ux-session-analysis
        description: "Searches session-transcripts, session-scenes across 1 feature index. Groups documents by participant_id."
        collections:
          - session-transcripts
          - session-scenes
        input_schema:
          query:
            type: text
            required: true
            description: "What to look for; searched across every index below"
        stages:
          - stage_name: search
            stage_id: feature_search
            parameters:
              searches:
                - feature_uri: "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1"
                  query:
                    input_mode: text
                    value: "{{INPUT.query}}"
                  top_k: 200
              fusion: rrf
              final_top_k: 200
          - stage_name: group
            stage_id: group_by
            parameters:
              group_by_field: participant_id
        tags:
          - template:ux-session-analysis
    Start building with Mixpeek

    Deploy this template, bring your data, and go from exploration to production.