Footage Intelligence at Archive Scale
Millions of scenes you can move through instead of search. Footage, ads and editor assets land in three clean collections, and scenes are clustered one partition at a time so every run stays under the sample-size caps.
A corpus you navigate by structure rather than by query, with every partition clustered under the limit that makes clustering possible at all.
Performance-video teams whose archive outgrew search. Once a corpus passes a million scenes, the question stops being 'find me this clip' and becomes 'show me what is in here', and a result list cannot answer that.
What deploys today7 ready6 need your inputWhat deploys today
What each part needs
The states are read from the manifest, so a part it ships commented out never shows as ready.
- raw-footage
- Deploys and processes documents as they arrive.
- footage-scenes
- Deploys and processes documents as they arrive.
- ad-creatives
- Deploys and processes documents as they arrive.
- editor-assets
- Deploys and processes documents as they arrive.
- scene-search
- Deploys and answers queries once its collections hold documents.
- ad-search
- Deploys and answers queries once its collections hold documents.
- scene-themes
- Deploys. Run it once per production_id by passing that filter each time you execute it; the manifest cannot hold the value.
- active-sync
- Connect this s3 source in Studio and sync it into its bucket. The manifest creates the bucket and leaves the connection to you.
- static-archive
- Connect this s3 source in Studio and sync it into its bucket. The manifest creates the bucket and leaves the connection to you.
- legacy-load
- Connect this s3 source in Studio and sync it into its bucket. The manifest creates the bucket and leaves the connection to you.
- ad-library
- Connect this s3 source in Studio and sync it into its bucket. The manifest creates the bucket and leaves the connection to you.
- editor-output
- You upload these files yourself. There is no connection to sync.
- cross-partition-themes
- The manifest cannot describe a cluster of clusters. Once the partitioned runs have completed, run it with POST /v1/clusters/{id}/execute in composite mode, naming at least two of those runs in source_execution_ids.
What it looks like
Pick a frame or a search and see what fires.
cleared- Ingest
- scene
- route: cleared
scenecleared
adcleared
assetclearedWhat you would ask it
Example searches this namespace answers once it is applied. Each one names the retriever that serves it.
“the take where the camera pushes in on the hands”
scene-search
scene segments across an archive nobody has watched end to end
“night exteriors with wet streets”
scene-search
returns the moment, with production_id carried through
“which cut did we use this shot in”
ad-search
the same footage matched against the delivered ads
How the namespace is wired
The diagram shows 5 buckets, 4 collections, 0 clean views and 2 retrievers. The manifest below applies 5 buckets, 4 collections (clean views included) and 2 retrievers today; the other parts are commented out in it, each with the reason. The diagram generates the manifest, so they cannot drift apart.
Reward signals
How reviewer decisions move the thresholds
Thresholds at ingest drift as the corpus changes. The reviewers working the queue are the ones who see where a threshold is wrong first, so this template routes their decisions back into the model that set it.
Labels stored with each cluster in the cluster's own collection, where a person can review and rename them. This version does not write labels back onto the scenes; the FAQ says why.
Which clusters get opened and searched from, which is the signal for whether the grouping is useful.
footage-scenesSystem collections in your namespace, on the same vector store as the rest of the template. They are yours to query.
New footage joins its partition's clusters on the next run of that partition, which re-cuts them when the shape has drifted.
One file spins up the namespace. Generated from the diagram above. Also served at /templates/footage-intelligence.namespace.yaml.
# footage-intelligence: one manifest spins up the namespace.
# Platform manifest schema (GET /v1/discovery/schema). Validate with POST /v1/manifest/validate,
# apply with POST /v1/manifest/apply or the Deploy button. Wiring comes from the flow diagram:
# edges are bucket -> collection sources, collection -> retriever scope, retriever -> view.
# Applying a SECOND time, over a namespace this template already created: use
# POST /v1/manifest/apply?mode=create_missing, which creates what is missing and leaves
# what exists alone. The default, create_only, fails the WHOLE apply and rolls it back if
# any resource already exists, so an upgrade looks like a dead end without this. Use
# mode=upsert to also patch resources that exist but have drifted from this file.
version: '1.0'
metadata:
name: footage-intelligence
description: "Namespace template footage-intelligence. Generated from the flow diagram on mixpeek.com/templates/footage-intelligence."
namespaces:
- name: footage-intelligence
description: "Everything below lives in this namespace."
feature_extractors:
- name: passthrough_extractor
version: v1
- name: multimodal_extractor
version: v1
# Daily spend budget, in dollars of what you are charged (1,000 credits a dollar).
# New batches pause once the namespace has spent this much today; queries are not capped.
# Raise it later on the namespace: PATCH /v1/namespaces/<id> {"spend_budget": {"daily_usd": N}}.
budget:
daily_usd: 0.1
# Data sources. A storage connection carries credentials, so it is created in Studio
# (or POST /v1/organizations/storage-connections) and synced into the bucket named here.
# active-sync: s3, continuous -> bucket footage-active
# static-archive: s3, one-time -> bucket footage-archive
# legacy-load: s3, one-time -> bucket footage-legacy
# ad-library: s3, continuous -> bucket ads
# editor-output: manual, on upload -> bucket editor-assets
buckets:
- name: footage-active
namespace: footage-intelligence
description: "Meant to be fed by active-sync (s3, continuous). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
schema:
properties:
content:
type: video
source_label:
type: string
filename:
type: string
production_id:
type: string
job_id:
type: string
- name: footage-archive
namespace: footage-intelligence
description: "Meant to be fed by static-archive (s3, one-time). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
schema:
properties:
content:
type: video
source_label:
type: string
filename:
type: string
production_id:
type: string
job_id:
type: string
- name: footage-legacy
namespace: footage-intelligence
description: "Meant to be fed by legacy-load (s3, one-time). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
schema:
properties:
content:
type: video
source_label:
type: string
filename:
type: string
production_id:
type: string
job_id:
type: string
- name: ads
namespace: footage-intelligence
description: "Meant to be fed by ad-library (s3, continuous). No source is connected yet: applying this manifest creates the bucket only. Connect the source to this bucket in Studio (Syncs) to start the feed."
schema:
properties:
content:
type: video
ad_id:
type: string
brand:
type: string
source_label:
type: string
- name: editor-assets
namespace: footage-intelligence
description: "Meant to be fed by editor-output (manual, on upload): upload files here or write them through the API. Nothing arrives until you do."
schema:
properties:
content:
type: video
asset_type:
type: string
source_label:
type: string
collections:
- name: raw-footage
namespace: footage-intelligence
description: "passthrough_extractor@v1 over bucket footage-active (the manifest wires one source bucket; footage-archive, footage-legacy are added after apply)."
source:
type: bucket
bucket: footage-active
feature_extractor:
name: passthrough_extractor
version: v1
field_passthrough:
- source_path: source_label
required: true
- source_path: filename
- source_path: production_id
- source_path: job_id
enabled: true
- name: footage-scenes
namespace: footage-intelligence
description: "multimodal_extractor@v1 over bucket footage-active (the manifest wires one source bucket; footage-archive, footage-legacy are added after apply). Feeds scene-search, scene-themes."
source:
type: bucket
bucket: footage-active
feature_extractor:
name: multimodal_extractor
version: v1
input_mappings:
video: content
field_passthrough:
- source_path: production_id
- source_path: job_id
- source_path: source_label
enabled: true
- name: ad-creatives
namespace: footage-intelligence
description: "multimodal_extractor@v1 over bucket ads. Feeds ad-search."
source:
type: bucket
bucket: ads
feature_extractor:
name: multimodal_extractor
version: v1
field_passthrough:
- source_path: ad_id
required: true
- source_path: brand
required: true
- source_path: source_label
enabled: true
- name: editor-assets
namespace: footage-intelligence
description: "passthrough_extractor@v1 over bucket editor-assets."
source:
type: bucket
bucket: editor-assets
feature_extractor:
name: passthrough_extractor
version: v1
field_passthrough:
- source_path: asset_type
required: true
- source_path: source_label
enabled: true
retrievers:
- name: scene-search
namespace: footage-intelligence
description: "Searches footage-scenes across 1 feature index."
collections:
- footage-scenes
input_schema:
query:
type: text
required: true
description: "What to look for; searched across every index below"
stages:
- stage_name: search
stage_id: feature_search
parameters:
searches:
- feature_uri: "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"
query:
input_mode: text
value: "{{INPUT.query}}"
top_k: 50
fusion: rrf
final_top_k: 50
tags:
- template:footage-intelligence
- name: ad-search
namespace: footage-intelligence
description: "Searches ad-creatives across 1 feature index."
collections:
- ad-creatives
input_schema:
query:
type: text
required: true
description: "What to look for; searched across every index below"
stages:
- stage_name: search
stage_id: feature_search
parameters:
searches:
- feature_uri: "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"
query:
input_mode: text
value: "{{INPUT.query}}"
top_k: 50
fusion: rrf
final_top_k: 50
tags:
- template:footage-intelligence
# Composite cluster run(s) are not emitted below, because the manifest schema cannot express them:
# cross-partition-themes: rolls up the centroids of scene-themes
# shared/manifest/models.py ClusterSpec requires source_collections to have at least one
# item, and a composite run has no source COLLECTION: it reads the centroids of prior
# EXECUTIONS. Naming it here fails validation for the entire manifest. Run it once the
# partitioned runs have completed: POST /v1/clusters/{id}/execute on the partitioned cluster
# with mode composite and source_execution_ids naming at least two of those runs.
clusters:
- name: scene-themes
namespace: footage-intelligence
description: "Groups footage-scenes by embedding similarity, one run per production_id."
source_collections:
- footage-scenes
cluster_type: vector
# Partitioned on production_id: execute once per value, passing
# filters: {AND: [{field: production_id, operator: eq, value: <one production_id>}]}
# with each execution. The manifest cannot hold that value.
# kmeans needs at least 24 documents in a production_id, since it has to
# put each one in a group. Raise or lower n_clusters to the number of groups you
# want to browse per production_id.
vector_config:
feature_uris:
- mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding
clustering_method: kmeans
algorithm_params:
n_clusters: 24
llm_labeling:
enabled: true
provider: google
model_name: gemini-2.5-flashDeploy this template, bring your data, and go from exploration to production.