NEWVectors or files. Pick a path.Start →
    Advertising
    Template v1.0 · updated 2026-09-14

    Contextual video match

    An article goes in, a ranked playlist of video comes out, and every match carries the entities, keywords and categories that produced it. Both sides are embedded twice, once by a text model and once by a multimodal one, so nothing is embedded at query time and a clip in another language still matches.

    Match video to an article by what the article is ABOUT, not by the words it shares. The hierarchy the buyer cares about is entity first, then keyword, then IAB category, and the retriever returns the overlap on each of the three beside every result, so an editor can read why a clip was chosen.

    Ad platforms and publishers placing video against editorial pages

    What deploys today8 ready2 need your input1 not available yet

    No marketplace starter set ships with this template: it applies empty and reads Your own video library and a handful of live article URLs, so connect your own data before querying it.

    287ms
    warm retrieval, no reranker
    Measured on the v8 benchmark of the pipeline this template is modelled on and recorded in that retriever's own header. Reranking was left out on purpose: a BGE reranker on CPU added about 30 seconds. A GPU reranker is a different trade and has not been measured here.
    122 / 118
    multimodal and text segments from 22 source videos
    From the reference deployment's benchmark record dated 2026-04-30: 22 video files between 13 and 149 MB produced 122 multimodal segments and 118 text segments. Two counts because the two extractors segment independently.
    233
    publisher URLs carrying page signals end to end
    The article side was run over 233 real publisher URLs, each returning entities with a salience score, keywords, an IAB path, sentiment with a confidence and a brand-safety flag. That is a coverage figure, not an accuracy figure.
    1408 + 1024
    dimensions per side, in two spaces
    mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding is 1408 dimensions and mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1 is 1024, both cosine. Verified against GET /v1/discovery/extractors on 2026-09-13.

    What deploys today

    video-scenesCollectionDeploys ready
    article-multimodalCollectionDeploys ready
    scene-transcriptsCollectionDeploys ready
    article-textCollectionDeploys ready
    contextual-matchRetrieverDeploys ready
    feature_searchRetriever stageDeploys ready
    deduplicateRetriever stageDeploys ready
    code_executionRetriever stageDeploys ready
    video-corpusData sourceNeeds your input
    article-feedData sourceNeeds your input
    attribute_filterRetriever stageNot available yet
    What each part needs

    The states are read from the manifest, so a part it ships commented out never shows as ready.

    video-scenes
    Deploys and processes documents as they arrive.
    article-multimodal
    Deploys and processes documents as they arrive.
    scene-transcripts
    Deploys and processes documents as they arrive.
    article-text
    Deploys and processes documents as they arrive.
    contextual-match
    Deploys and answers queries once its collections hold documents.
    feature_search
    The one-click manifest includes this stage, and every search runs from a document query.
    deduplicate
    The one-click manifest includes this stage.
    code_execution
    The one-click manifest includes this stage.
    video-corpus
    Create the s3 credentials as manifest secrets, then uncomment the sync. It ships commented out because apply tests a connection as it creates it.
    article-feed
    You upload these files yourself. There is no connection to sync.
    attribute_filter
    The one-click manifest has no attribute_filter stage. Add it to the retriever after deploy.

    What it looks like

    Pick a frame or a search and see what fires.

    Simulated walkthrough · illustrative frames, scripted decisions
    A video editing timeline with clip thumbnails on the video track and an audio waveform belowcleared
    video-scenes -> scene-transcriptsStill: Pexels / Alex Fuframe 1 of 1
    Detections on this frame
    Decision path
    1. video-scenes -> scene-transcripts
    2. scene segments
    3. transcript
    4. route: cleared
    each scene embedded separately, matched on its own
    Running tally
    1
    cleared
    0
    review
    0
    excluded
    What a person's call does here
    None. No editor thumbs, no accepted-playlist marks are recorded by the template as shipped.
    Try a search
    Pick a search to see which retriever answers it and whether it works on a one-click deploy.
    Where the decisions fire
    Illustrative frames; the boxes are authored to show the decision path
    A video editing timeline with clip thumbnails on the video track and an audio waveform belowscene segmentstranscriptcleared
    video-scenes -> scene-transcripts · The file is cut where the scene changes and each scene is embedded on its own, so an article matches the forty seconds that fit it.Still: Pexels / Alex Fu
    What goes in
    Connect your data to get started.
    videos
    One object per video. The template cuts each into segments and indexes transcription, on-screen text and a multimodal embedding per segment.
    articles
    One object per page: the article text, plus article_url, published_at and tenant_id as fields. The tenant field is what a query-time scope filter reads.
    What comes out
    Outputs from this template.
    video-scenes
    Deploys ready
    Segment documents with a 1408-dimension multimodal embedding, timing, transcription, OCR text, and the entities, keywords and IAB categories read off the segment.
    scene-transcripts
    Deploys ready
    The spoken side of each segment as a 1024-dimension text embedding, so a text-first query reaches it directly.
    article-multimodal
    Deploys ready
    The article in the same 1408-dimension space as the video, which is what makes the cross-modal half of the match a document reference rather than a query embedding.
    article-text
    Deploys ready
    The article as a 1024-dimension text embedding, with its URL, publish time and tenant carried as fields.

    What you would ask it

    Example searches this namespace answers once it is applied. Each one names the retriever that serves it.

    • “a kitchen scene with someone cooking at a hob”

      contextual-match

      scene segments ranked by how well the frame and the words fit the query

    • “commuters on a crowded platform at rush hour”

      contextual-match

      returns the seconds that match, so a placement lands inside them

    • “quiet interior, one person, natural light”

      contextual-match

      the scene embedding carries composition, so mood-shaped queries work

    How the namespace is wired

    The diagram shows 2 buckets, 4 collections, 0 clean views and 1 retrievers. The manifest below applies 2 buckets, 4 collections (clean views included) and 1 retrievers today; the other parts are commented out in it, each with the reason. The diagram generates the manifest, so they cannot drift apart.

    SourceBucketCollectionClean viewRetrieverClusterConnectionBucket syncAlertTriggerClick a node to inspect it
    syncsyncsearchsearchsearchsearch

    Reward signals

    How reviewer decisions move the thresholds

    Thresholds at ingest drift as the corpus changes. The reviewers working the queue are the ones who see where a threshold is wrong first, so this template routes their decisions back into the model that set it.

    Nothing in this template writes interaction signals back. Saying so is better than implying a flywheel that is not wired, and the ranking is deterministic given the same corpus.
    Explicit signals

    None. No editor thumbs, no accepted-playlist marks are recorded by the template as shipped.

    Implicit signals

    None from the application. Mixpeek records retriever executions server-side, which is telemetry about queries rather than feedback about which match was right.

    Where they land
    mxp_retriever_executions

    System collections in your namespace, on the same vector store as the rest of the template. They are yours to query.

    How the loop closes

    The loop that would close it is the customer's own answer key: a set of articles with the playlist they consider correct. That is an evaluation input rather than a signal the template can collect on its own.

    One file spins up the namespace. Generated from the diagram above. Also served at /templates/contextual-video-match.namespace.yaml.

    # contextual-video-match: one manifest spins up the namespace.
    # Platform manifest schema (GET /v1/discovery/schema). Validate with POST /v1/manifest/validate,
    # apply with POST /v1/manifest/apply or the Deploy button. Wiring comes from the flow diagram:
    # edges are bucket -> collection sources, collection -> retriever scope, retriever -> view.
    # Applying a SECOND time, over a namespace this template already created: use
    # POST /v1/manifest/apply?mode=create_missing, which creates what is missing and leaves
    # what exists alone. The default, create_only, fails the WHOLE apply and rolls it back if
    # any resource already exists, so an upgrade looks like a dead end without this. Use
    # mode=upsert to also patch resources that exist but have drifted from this file.
    version: '1.0'
    metadata:
      name: contextual-video-match
      description: "Namespace template contextual-video-match. Generated from the flow diagram on mixpeek.com/templates/contextual-video-match."
    namespaces:
      - name: contextual-video-match
        description: "Everything below lives in this namespace."
        feature_extractors:
          - name: multimodal_extractor
            version: v1
          - name: text_extractor
            version: v1
        # Daily spend budget, in dollars of what you are charged (1,000 credits a dollar).
        # New batches pause once the namespace has spent this much today; queries are not capped.
        # Raise it later on the namespace: PATCH /v1/namespaces/<id> {"spend_budget": {"daily_usd": N}}.
        budget:
          daily_usd: 0.1
    
    # Data sources. A storage connection carries credentials, so it is created in Studio
    # (or POST /v1/organizations/storage-connections) and synced into the bucket named here.
    #   video-corpus: s3, continuous, the video library you want matched against, in your account -> bucket videos
    #   article-feed: manual, continuous, one object per page: the article text, its URL, its publish time and the tenant it belongs to -> bucket articles
    buckets:
      - name: videos
        namespace: contextual-video-match
        description: "Meant to be fed by video-corpus (s3, continuous). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
        schema:
          properties:
            content:
              type: video
      - name: articles
        namespace: contextual-video-match
        description: "Meant to be fed by article-feed (manual, continuous): upload files here or write them through the API. Nothing arrives until you do."
        schema:
          properties:
            content:
              type: text
    
    # Connect your own storage. Create the two secrets, uncomment, and apply again;
    # apply live-tests the connection, so it must have real credentials to succeed.
    # Until then the buckets above accept direct uploads.
    # storage_connections:
    #   - name: video-corpus-connection
    #     provider: s3
    #     description: "Read-only access to the s3 location holding the video library you want matched against, in your account."
    #     config:
    #       region: us-east-1
    #       credentials:
    #         type: access_key
    #         access_key_id: ${{ secrets.AWS_ACCESS_KEY_ID }}
    #         secret_access_key: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
    
    # bucket_syncs:
    #   - name: video-corpus-sync
    #     bucket: videos
    #     connection: video-corpus-connection
    #     source_path: "videos/"
    #     sync_mode: continuous
    #     polling_interval_seconds: 300
    #     skip_duplicates: true
    #     file_filters:
    #       include_patterns: ["*.mp4", "*.mov", "*.webm"]
    #     schema_mapping:
    #       mappings:
    #         content:
    #           target_type: blob
    #           source: {type: file}
    #           blob_type: video
    
    collections:
      - name: video-scenes
        namespace: contextual-video-match
        description: "multimodal_extractor@v1 over bucket videos. Feeds scene-transcripts, contextual-match."
        source:
          type: bucket
          bucket: videos
        feature_extractor:
          name: multimodal_extractor
          version: v1
          parameters:
            run_transcription: true
            run_video_description: true
            run_ocr: true
            description_prompt: "Identify the named entities and topical keywords present in this video segment, from what is spoken, shown on screen, and visible in the frame. Transcribe any text burned into the frame, such as captions, lower thirds, tickers and signage, verbatim."
            response_shape:
              type: "object"
              properties:
                entities:
                  type: "array"
                  items:
                    type: "string"
                  description: "Named entities (people, organizations, places) in canonical English form, deduplicated. 3-10 items."
                keywords:
                  type: "array"
                  items:
                    type: "string"
                  description: "Topical keywords or short phrases describing what the segment is about. 3-8 items."
                iab_categories:
                  type: "array"
                  items:
                    type: "string"
                  description: "IAB Content Taxonomy tier-1 category names that describe this segment, written exactly as IAB names them, for example \"News & Politics\", \"Arts & Entertainment\", \"Sports\", \"Automotive\". 1-3 items."
                ocr_text:
                  type: "string"
                  description: "Text burned into the frame, transcribed verbatim: captions, lower thirds, tickers, signage. Empty string when the segment shows no text."
          input_mappings:
            video: content
          field_passthrough:
            - source_path: file_location
            - source_path: segment_id
            - source_path: start_time
            - source_path: end_time
            - source_path: tenant_id
        enabled: true
      - name: article-multimodal
        namespace: contextual-video-match
        description: "multimodal_extractor@v1 over bucket articles. Feeds contextual-match."
        source:
          type: bucket
          bucket: articles
        feature_extractor:
          name: multimodal_extractor
          version: v1
          input_mappings:
            text: content
          field_passthrough:
            - source_path: article_url
            - source_path: published_at
            - source_path: tenant_id
        enabled: true
      - name: scene-transcripts
        namespace: contextual-video-match
        description: "text_extractor@v1 over collection video-scenes. Feeds contextual-match."
        source:
          type: collection
          collection: video-scenes
        feature_extractor:
          name: text_extractor
          version: v1
          input_mappings:
            text: transcription
          field_passthrough:
            - source_path: segment_id
            - source_path: video_segment_url
            - source_path: start_time
            - source_path: end_time
        enabled: true
      - name: article-text
        namespace: contextual-video-match
        description: "text_extractor@v1 over bucket articles. Feeds contextual-match."
        source:
          type: bucket
          bucket: articles
        feature_extractor:
          name: text_extractor
          version: v1
          input_mappings:
            text: content
          field_passthrough:
            - source_path: article_url
            - source_path: published_at
            - source_path: tenant_id
        enabled: true
    retrievers:
      - name: contextual-match
        namespace: contextual-video-match
        description: "Matches scene-transcripts and video-scenes against a document from article-text and article-multimodal, with no query-time embedding."
        collections:
          - scene-transcripts
          - video-scenes
          - article-text
          - article-multimodal
        input_schema:
          article_text_doc_id:
            type: string
            required: true
            description: "Document id of the article in article-text, whose stored text vector the scene transcripts are matched against"
          article_multimodal_doc_id:
            type: string
            required: true
            description: "Document id of the same article in article-multimodal, whose stored multimodal vector the video scenes are matched against"
          article_entities:
            type: array
            required: false
            description: "Entities already extracted from the article; drives the 0.25 overlap term"
          article_keywords:
            type: array
            required: false
            description: "Keywords already extracted from the article; drives the 0.15 overlap term"
          article_iab_categories:
            type: array
            required: false
            description: "IAB categories for the article; drives the 0.10 taxonomy term and the Entertainment-against-News penalty"
        stages:
          - stage_name: search
            stage_id: feature_search
            parameters:
              searches:
                - feature_uri: "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1"
                  collection_identifiers:
                    - scene-transcripts
                  query:
                    input_mode: document
                    document_ref:
                      collection_id: article-text
                      document_id: "{{INPUT.article_text_doc_id}}"
                  top_k: 30
                  min_score: 0.15
                - feature_uri: "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"
                  collection_identifiers:
                    - video-scenes
                  query:
                    input_mode: document
                    document_ref:
                      collection_id: article-multimodal
                      document_id: "{{INPUT.article_multimodal_doc_id}}"
                  top_k: 30
              fusion: rrf
              final_top_k: 30
          - stage_name: deduplicate
            stage_id: deduplicate
            parameters:
              strategy: field
              fields:
                - source_video_url
                - start_time
              keep: first
          - stage_name: score
            stage_id: code_execution
            parameters:
              language: python
              output_field: scoring
              timeout_ms: 5000
              code: |
                _ae = {{ INPUT.article_entities }}
                _ak = {{ INPUT.article_keywords }}
                _ac = {{ INPUT.article_iab_categories }}
                art_ents = set(str(e).lower() for e in (_ae or []))
                art_kws = set(str(k).lower() for k in (_ak or []))
                art_cats = set(str(c).lower() for c in (_ac or []))
                result = []
                for doc in docs:
                    ve = set(str(e).lower() for e in doc.get('entities', []))
                    vk = set(str(k).lower() for k in doc.get('keywords', []))
                    vc = set(str(c).lower() for c in doc.get('iab_categories', []))
                    vt = ((doc.get('transcription') or '') + ' ' + (doc.get('ocr_text') or '')).lower()
                    te = set(e for e in art_ents if e in vt)
                    all_ents = (art_ents & ve) | te
                    eo = len(all_ents) / max(len(art_ents), 1) if art_ents else 0
                    ko = len(art_kws & vk) / max(len(art_kws), 1) if art_kws else 0
                    to = len(art_cats & vc) / max(len(art_cats), 1) if art_cats else 0
                    s = doc.get('scores', {}).get('search', doc.get('score', 0))
                    cs = 0.50 * s + 0.25 * min(eo * 1.5, 1.0) + 0.15 * min(ko * 1.5, 1.0) + 0.10 * to
                    tp = 0.7 if vc and vc & {'entertainment', 'arts & entertainment'} and art_cats & {'news & politics', 'law government & politics'} else 1.0
                    cs *= tp
                    result.append({'contextual_score': round(cs, 4), 'entity_overlap': round(eo, 4), 'keyword_overlap': round(ko, 4), 'taxonomy_overlap': round(to, 4), 'entities_matched': list(all_ents)[:5], 'keywords_matched': list(art_kws & vk)[:5], 'iab_matched': list(art_cats & vc)[:3], 'taxonomy_penalty': tp})
        tags:
          - template:contextual-video-match
    Start building with Mixpeek

    Deploy this template, bring your data, and go from exploration to production.