NEWVectors or files. Pick a path.Start →
    Search & Discovery
    7 min read
    Updated 2026-10-07

    How Do I Match the Right Video to Each Article on My Site?

    To match the right video to each article, index videos as short scenes with transcript, on-screen text and an embedding, extract each article's entities, keywords and IAB category, and rank scenes by shared entities first. Return the matching seconds and why each was chosen.

    Contextual Advertising
    Video Recommendations
    Publishers
    IAB Taxonomy
    Video Search

    How do I match the right video to each article on my site?



    Match on what the article is about. Index your video library in short scenes, each with its transcript, the text shown on screen and a visual embedding. For every article, extract the people, companies and places it names, its keywords and its IAB category, and embed the article too. Then rank scenes by how much they share with the article: entities first, keywords next, category last, with embedding similarity to break ties. Return the matching seconds of each video with the reasons it was chosen, so an editor or an ad server can check the pick.

    Why do simple approaches pick the wrong video?



    ApproachWhat goes wrong
    Matching article tags to video tagsTags are sparse and written by different people, so most articles match nothing or match the same few videos
    Matching the article title to video titlesA video titled "Weekend highlights" can cover the exact story; the title never says so
    Embedding the whole video onceOne vector for a 20-minute video blurs every topic in it, and the result cannot say where the relevant part is
    Picking by IAB category aloneHundreds of articles share a category, so every one of them gets the same video

    What should the matching signals be?



    1. Entities. Named people, organisations, products and places, read from the article text and from each scene's transcript and on-screen text. Sharing a named entity is the strongest evidence two things are about the same story. 2. Keywords. The distinctive terms of the article, matched against the scene's words. 3. Category. The IAB content taxonomy path, which keeps a match in the right neighbourhood and is what ad buyers target. 4. Embedding similarity. The article and the scene in the same vector space, which catches matches with no shared words, including clips in another language.

    On-screen text matters more than it looks. A clip with a garbled or foreign-language transcript usually still shows a name or a caption in the frame, and reading it is what keeps that clip matchable. Why video search misses on-screen text covers how.

    How fast does it need to be?



    Fast enough to run while the page loads. Do the expensive work ahead of time: index every video and every article when they are published, so a page view runs one search over stored vectors and embeds nothing. Measured on the reference pipeline behind our template, warm retrieval took 287 ms without a reranker. A reranker running on CPU added about 30 seconds, which is why it was left out of the request path.

    How do I show why a video was chosen?



    Return the evidence with every match: the entities, keywords and categories the article and the scene share, and each signal's contribution to the score. Editors can then see that a clip was placed because both name the same company, and fix a bad rule directly. Mixpeek's design notes list "no trace of why a result came back" among the failures that sink systems like this one. A match that carries its own reasons avoids that.

    Brand safety is a separate check. Run it on both sides before a placement goes live; brand safety vs brand suitability explains the difference.

    What results should I expect?



    From the reference deployment of the pipeline our template is built on:

  1. 22 source videos between 13 and 149 MB produced 122 multimodal segments and 118 text segments.
  2. 233 real publisher URLs went through the article side, each returning entities with a salience score,
  3. keywords, an IAB path, sentiment and a brand-safety flag. That figure measures coverage; accuracy was not measured.
  4. 287 ms warm retrieval per article with nothing embedded at request time.


  5. How do I do this with Mixpeek?



    Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. The Contextual video match template deploys the pipeline above: videos cut into scenes with transcription, on-screen text and a multimodal embedding; articles with their entities, keywords and IAB categories; and a retriever that takes an article already in the index and returns ranked video segments with start and end times and the overlap on each signal. Video indexing starts at $0.05 a minute on the rate card.

    It suits publishers and ad platforms placing video against editorial pages at volume. A site with a few dozen videos can choose by hand. Related: how to find every ad that shows a product and AI platforms for ad tech.

    Frequently Asked Questions



    How do I match the right video to each article on my site?



    Index your videos as short scenes with transcript, on-screen text and a visual embedding, extract each article's entities, keywords and IAB category, and rank scenes by shared entities first, then keywords, then category, with embedding similarity as the tie-breaker. Return the matching seconds and the reasons.

    Can it match a video in another language to an English article?



    Yes, through the embedding and the on-screen text. Embeddings place related content close together across languages, and names shown in the frame match even when the speech does not.

    Does contextual matching need cookies or user data?



    No. It reads only the page and the video, so it works without third-party cookies or a user profile.

    How do I stop the same video appearing on every article?



    Rank on entities and keywords before category, and cap how often one video can be placed per day. Category alone gives hundreds of articles the same top match.
    Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs