NEWVectors or files. Pick a path.Start →
    Media
    Template v1.2 · updated 2026-09-15

    Footage Intelligence at Archive Scale

    Millions of scenes you can move through instead of search. Footage, ads and editor assets land in three clean collections, and scenes are clustered one partition at a time so every run stays under the sample-size caps.

    A corpus you navigate by structure rather than by query, with every partition clustered under the limit that makes clustering possible at all.

    Performance-video teams whose archive outgrew search. Once a corpus passes a million scenes, the question stops being 'find me this clip' and becomes 'show me what is in here', and a result list cannot answer that.

    What deploys today7 ready6 need your input
    1.27M
    scenes in the reference corpus
    The scale the partitioning recipe was written against, from the blessed clustering runbook.
    5 → 3
    sources into clean collections
    Three footage sources merge into one collection; ads and editor assets keep their own.
    100k
    documents per HDBSCAN run
    The hard cap on the quadratic methods. The pairwise distance matrix costs sample_size squared times 8 bytes, about 80 GB at that ceiling.
    1M
    documents per k-means run
    The hard cap on the linear methods. Above it, a partition is split again rather than the cap being raised.

    What deploys today

    raw-footageCollectionDeploys ready
    footage-scenesCollectionDeploys ready
    ad-creativesCollectionDeploys ready
    editor-assetsCollectionDeploys ready
    scene-searchRetrieverDeploys ready
    ad-searchRetrieverDeploys ready
    scene-themesClusterDeploys ready
    active-syncData sourceNeeds your input
    static-archiveData sourceNeeds your input
    legacy-loadData sourceNeeds your input
    ad-libraryData sourceNeeds your input
    editor-outputData sourceNeeds your input
    cross-partition-themesClusterNeeds your input
    What each part needs

    The states are read from the manifest, so a part it ships commented out never shows as ready.

    raw-footage
    Deploys and processes documents as they arrive.
    footage-scenes
    Deploys and processes documents as they arrive.
    ad-creatives
    Deploys and processes documents as they arrive.
    editor-assets
    Deploys and processes documents as they arrive.
    scene-search
    Deploys and answers queries once its collections hold documents.
    ad-search
    Deploys and answers queries once its collections hold documents.
    scene-themes
    Deploys. Run it once per production_id by passing that filter each time you execute it; the manifest cannot hold the value.
    active-sync
    Connect this s3 source in Studio and sync it into its bucket. The manifest creates the bucket and leaves the connection to you.
    static-archive
    Connect this s3 source in Studio and sync it into its bucket. The manifest creates the bucket and leaves the connection to you.
    legacy-load
    Connect this s3 source in Studio and sync it into its bucket. The manifest creates the bucket and leaves the connection to you.
    ad-library
    Connect this s3 source in Studio and sync it into its bucket. The manifest creates the bucket and leaves the connection to you.
    editor-output
    You upload these files yourself. There is no connection to sync.
    cross-partition-themes
    The manifest cannot describe a cluster of clusters. Once the partitioned runs have completed, run it with POST /v1/clusters/{id}/execute in composite mode, naming at least two of those runs in source_execution_ids.

    What it looks like

    Pick a frame or a search and see what fires.

    Simulated walkthrough · illustrative frames, scripted decisions
    A camera operator silhouetted against a lit setcleared
    IngestStill: Pexels / Wolrider YURTSEVENframe 1 of 3
    Detections on this frame
    Decision path
    1. Ingest
    2. scene
    3. route: cleared
    Landed by a passthrough extractor with its source label intact.
    Running tally
    1
    cleared
    0
    review
    0
    excluded
    What a person's call does here
    Labels stored with each cluster in the cluster's own collection, where a person can review and rename them. This version does not write labels back onto the scenes; the FAQ says why.
    Try a search
    Pick a search to see which retriever answers it and whether it works on a one-click deploy.
    Where the decisions fire
    Illustrative frames; the boxes are authored to show the decision path
    A camera operator silhouetted against a lit setscenecleared
    Ingest · Raw footage arrives from three sources that share one schema, so the merge into a single collection is homogeneous.Still: Pexels / Wolrider YURTSEVEN
    A phone on a gimbal recording against a green screenadcleared
    Ingest · Ads keep their own identifiers, so a cluster can be read back to the creative it came from.Still: Pexels / Javier Gonzalez
    Two editors reviewing a timeline on a studio monitorassetcleared
    Ingest · Editor assets land in their own collection rather than being mixed into footage, because they are a different kind of thing.Still: Pexels / Ron Lach
    What goes in
    Connect your data to get started.
    Footage
    Three sources: an active sync, a large static archive, and a frozen legacy load. One shared schema so the merge is homogeneous.
    Ads
    Approved creatives, carrying their own ad and brand identifiers.
    Editor assets
    Graphics, text capsules and generated assets from the edit bay.
    What comes out
    Outputs from this template.
    Clean collections
    Raw footage, ads and editor assets, landed by a passthrough extractor so downstream teams build on stable inputs.
    Per-partition clusters
    One clustering execution per partition key, written to the cluster's own collection. Each run's cluster ids start again at cl_0, so a cluster id identifies a group within its run rather than across the corpus.
    A composite roll-up
    The centroids of the partition runs you name, grouped into a layer that sits above them in the same collection.
    Filtered retrievers
    Scene and ad search that can be narrowed to one partition.

    What you would ask it

    Example searches this namespace answers once it is applied. Each one names the retriever that serves it.

    • “the take where the camera pushes in on the hands”

      scene-search

      scene segments across an archive nobody has watched end to end

    • “night exteriors with wet streets”

      scene-search

      returns the moment, with production_id carried through

    • “which cut did we use this shot in”

      ad-search

      the same footage matched against the delivered ads

    How the namespace is wired

    The diagram shows 5 buckets, 4 collections, 0 clean views and 2 retrievers. The manifest below applies 5 buckets, 4 collections (clean views included) and 2 retrievers today; the other parts are commented out in it, each with the reason. The diagram generates the manifest, so they cannot drift apart.

    SourceBucketCollectionClean viewRetrieverClusterConnectionBucket syncAlertTriggerClick a node to inspect it
    syncsyncsyncsyncsyncsearchsearchclusterroll up

    Reward signals

    How reviewer decisions move the thresholds

    Thresholds at ingest drift as the corpus changes. The reviewers working the queue are the ones who see where a threshold is wrong first, so this template routes their decisions back into the model that set it.

    Clusters are re-cut as the corpus grows, one partition at a time.
    Explicit signals

    Labels stored with each cluster in the cluster's own collection, where a person can review and rename them. This version does not write labels back onto the scenes; the FAQ says why.

    Implicit signals

    Which clusters get opened and searched from, which is the signal for whether the grouping is useful.

    Where they land
    footage-scenes

    System collections in your namespace, on the same vector store as the rest of the template. They are yours to query.

    How the loop closes

    New footage joins its partition's clusters on the next run of that partition, which re-cuts them when the shape has drifted.

    One file spins up the namespace. Generated from the diagram above. Also served at /templates/footage-intelligence.namespace.yaml.

    # footage-intelligence: one manifest spins up the namespace.
    # Platform manifest schema (GET /v1/discovery/schema). Validate with POST /v1/manifest/validate,
    # apply with POST /v1/manifest/apply or the Deploy button. Wiring comes from the flow diagram:
    # edges are bucket -> collection sources, collection -> retriever scope, retriever -> view.
    # Applying a SECOND time, over a namespace this template already created: use
    # POST /v1/manifest/apply?mode=create_missing, which creates what is missing and leaves
    # what exists alone. The default, create_only, fails the WHOLE apply and rolls it back if
    # any resource already exists, so an upgrade looks like a dead end without this. Use
    # mode=upsert to also patch resources that exist but have drifted from this file.
    version: '1.0'
    metadata:
      name: footage-intelligence
      description: "Namespace template footage-intelligence. Generated from the flow diagram on mixpeek.com/templates/footage-intelligence."
    namespaces:
      - name: footage-intelligence
        description: "Everything below lives in this namespace."
        feature_extractors:
          - name: passthrough_extractor
            version: v1
          - name: multimodal_extractor
            version: v1
        # Daily spend budget, in dollars of what you are charged (1,000 credits a dollar).
        # New batches pause once the namespace has spent this much today; queries are not capped.
        # Raise it later on the namespace: PATCH /v1/namespaces/<id> {"spend_budget": {"daily_usd": N}}.
        budget:
          daily_usd: 0.1
    
    # Data sources. A storage connection carries credentials, so it is created in Studio
    # (or POST /v1/organizations/storage-connections) and synced into the bucket named here.
    #   active-sync: s3, continuous -> bucket footage-active
    #   static-archive: s3, one-time -> bucket footage-archive
    #   legacy-load: s3, one-time -> bucket footage-legacy
    #   ad-library: s3, continuous -> bucket ads
    #   editor-output: manual, on upload -> bucket editor-assets
    buckets:
      - name: footage-active
        namespace: footage-intelligence
        description: "Meant to be fed by active-sync (s3, continuous). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
        schema:
          properties:
            content:
              type: video
            source_label:
              type: string
            filename:
              type: string
            production_id:
              type: string
            job_id:
              type: string
      - name: footage-archive
        namespace: footage-intelligence
        description: "Meant to be fed by static-archive (s3, one-time). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
        schema:
          properties:
            content:
              type: video
            source_label:
              type: string
            filename:
              type: string
            production_id:
              type: string
            job_id:
              type: string
      - name: footage-legacy
        namespace: footage-intelligence
        description: "Meant to be fed by legacy-load (s3, one-time). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
        schema:
          properties:
            content:
              type: video
            source_label:
              type: string
            filename:
              type: string
            production_id:
              type: string
            job_id:
              type: string
      - name: ads
        namespace: footage-intelligence
        description: "Meant to be fed by ad-library (s3, continuous). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
        schema:
          properties:
            content:
              type: video
            ad_id:
              type: string
            brand:
              type: string
            source_label:
              type: string
      - name: editor-assets
        namespace: footage-intelligence
        description: "Meant to be fed by editor-output (manual, on upload): upload files here or write them through the API. Nothing arrives until you do."
        schema:
          properties:
            content:
              type: video
            asset_type:
              type: string
            source_label:
              type: string
    collections:
      - name: raw-footage
        namespace: footage-intelligence
        description: "passthrough_extractor@v1 over bucket footage-active (the manifest wires one source bucket; footage-archive, footage-legacy are added after apply)."
        source:
          type: bucket
          bucket: footage-active
        feature_extractor:
          name: passthrough_extractor
          version: v1
          field_passthrough:
            - source_path: source_label
              required: true
            - source_path: filename
            - source_path: production_id
            - source_path: job_id
        enabled: true
      - name: footage-scenes
        namespace: footage-intelligence
        description: "multimodal_extractor@v1 over bucket footage-active (the manifest wires one source bucket; footage-archive, footage-legacy are added after apply). Feeds scene-search, scene-themes."
        source:
          type: bucket
          bucket: footage-active
        feature_extractor:
          name: multimodal_extractor
          version: v1
          input_mappings:
            video: content
          field_passthrough:
            - source_path: production_id
            - source_path: job_id
            - source_path: source_label
        enabled: true
      - name: ad-creatives
        namespace: footage-intelligence
        description: "multimodal_extractor@v1 over bucket ads. Feeds ad-search."
        source:
          type: bucket
          bucket: ads
        feature_extractor:
          name: multimodal_extractor
          version: v1
          field_passthrough:
            - source_path: ad_id
              required: true
            - source_path: brand
              required: true
            - source_path: source_label
        enabled: true
      - name: editor-assets
        namespace: footage-intelligence
        description: "passthrough_extractor@v1 over bucket editor-assets."
        source:
          type: bucket
          bucket: editor-assets
        feature_extractor:
          name: passthrough_extractor
          version: v1
          field_passthrough:
            - source_path: asset_type
              required: true
            - source_path: source_label
        enabled: true
    retrievers:
      - name: scene-search
        namespace: footage-intelligence
        description: "Searches footage-scenes across 1 feature index."
        collections:
          - footage-scenes
        input_schema:
          query:
            type: text
            required: true
            description: "What to look for; searched across every index below"
        stages:
          - stage_name: search
            stage_id: feature_search
            parameters:
              searches:
                - feature_uri: "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"
                  query:
                    input_mode: text
                    value: "{{INPUT.query}}"
                  top_k: 50
              fusion: rrf
              final_top_k: 50
        tags:
          - template:footage-intelligence
      - name: ad-search
        namespace: footage-intelligence
        description: "Searches ad-creatives across 1 feature index."
        collections:
          - ad-creatives
        input_schema:
          query:
            type: text
            required: true
            description: "What to look for; searched across every index below"
        stages:
          - stage_name: search
            stage_id: feature_search
            parameters:
              searches:
                - feature_uri: "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"
                  query:
                    input_mode: text
                    value: "{{INPUT.query}}"
                  top_k: 50
              fusion: rrf
              final_top_k: 50
        tags:
          - template:footage-intelligence
    # Composite cluster run(s) are not emitted below, because the manifest schema cannot express them:
    #   cross-partition-themes: rolls up the centroids of scene-themes
    # shared/manifest/models.py ClusterSpec requires source_collections to have at least one
    # item, and a composite run has no source COLLECTION: it reads the centroids of prior
    # EXECUTIONS. Naming it here fails validation for the entire manifest. Run it once the
    # partitioned runs have completed: POST /v1/clusters/{id}/execute on the partitioned cluster
    # with mode composite and source_execution_ids naming at least two of those runs.
    clusters:
      - name: scene-themes
        namespace: footage-intelligence
        description: "Groups footage-scenes by embedding similarity, one run per production_id."
        source_collections:
          - footage-scenes
        cluster_type: vector
        # Partitioned on production_id: execute once per value, passing
        #   filters: {AND: [{field: production_id, operator: eq, value: <one production_id>}]}
        # with each execution. The manifest cannot hold that value.
        # kmeans needs at least 24 documents in a production_id, since it has to
        # put each one in a group. Raise or lower n_clusters to the number of groups you
        # want to browse per production_id.
        vector_config:
          feature_uris:
            - mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding
          clustering_method: kmeans
          algorithm_params:
            n_clusters: 24
        llm_labeling:
          enabled: true
          provider: google
          model_name: gemini-2.5-flash
    Start building with Mixpeek

    Deploy this template, bring your data, and go from exploration to production.