Automated Video Tagging
Automatically generate descriptive tags for video content using scene analysis, object detection, and taxonomy classification. Each video receives structured labels for scenes, objects, actions, and custom business categories without manual annotation.
import requestsfrom mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="video-tags")API = "https://api.mixpeek.com/v1"HEADERS = {"Authorization": "Bearer YOUR_API_KEY", "X-Namespace": "video-tags"}# 1. Category exemplars: short clips labeled sports, cooking, technology, outdoorsexample_bucket = client.buckets.create(bucket_name="tag-examples",bucket_schema={"properties": {"clip": {"type": "video"}}},)examples = client.collections.create(collection_name="tag-examples",source={"type": "bucket", "bucket_ids": [example_bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1",},)client.buckets.upload(example_bucket["bucket_id"],blobs=[{"property": "clip", "type": "video", "data": "s3://your-bucket/tags/cooking-01.mp4"}],)client.collections.trigger(examples["collection_id"])# 2. The retriever the taxonomy uses to find the closest exemplarmatcher = client.retrievers.create(retriever_name="tag-matcher",collection_identifiers=["tag-examples"],input_schema={"query": {"type": "text", "required": True}},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding","query": {"input_mode": "text", "value": "{{INPUT.query}}"},"top_k": 3,},],"final_top_k": 3,},},],)# A flat taxonomy matches each document, through a retriever, against a# collection of labeled examples. The SDK has no taxonomies resource, so this is REST.taxonomy = requests.post(API + "/taxonomies", headers=HEADERS, json={"taxonomy_name": "video_tags","config": {"taxonomy_type": "flat","retriever_id": matcher["retriever_id"],"input_mappings": [{"input_key": "query", "source_type": "payload", "path": "description"}],"source_collection": {"collection_id": examples["collection_id"]},},}).json()# 3. The video collection materializes the tags as scenes are writtenbucket = client.buckets.create(bucket_name="videos",bucket_schema={"properties": {"video": {"type": "video"}}},)videos = client.collections.create(collection_name="videos",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"split_method": "scene","run_video_description": True,},},taxonomy_applications=[{"taxonomy_id": taxonomy["taxonomy_id"], "execution_mode": "materialize"}],)client.buckets.upload(bucket["bucket_id"],blobs=[{"property": "video", "type": "video", "data": "s3://your-bucket/videos/weekend-hike.mp4"}],)client.collections.trigger(videos["collection_id"])# Videos indexed before the taxonomy was attached are labeled with apply-taxonomyrequests.post(API + "/collections/videos/apply-taxonomy", headers=HEADERS, json={"taxonomy_id": taxonomy["taxonomy_id"]})
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
taxonomy enrich
Classify documents against taxonomy nodes via vector similarity
Use Cases Using This Recipe
Sports Highlights
Auto-generate highlight reels from full-length sports footage
24x faster
Highlight generation time
Sports broadcasters, media companies, and content teams processing 100+ hours of live footage weekly
Media Archive Intelligence
Transform decades of media archives into searchable, monetizable assets
12x more content findable
Archive discoverability
Broadcasters, news organizations, film studios, music labels, and cultural institutions managing archives of 10,000+ hours of video, millions of photographs, or extensive audio collections
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Hierarchical Classification
Assign content to multi-level category hierarchies using embedding-based classification. Define your taxonomy once, then classify new content automatically with confidence scores.
Brand Safety & Ad Verification Pipeline
GARM-compliant brand safety pipeline for ad networks. Analyze video and image creatives for brand safety violations before serving.
Video Scene Search
Find specific scenes within videos using natural language descriptions. The pipeline detects scene boundaries, generates embeddings for each scene, and enables precise timestamp-level search across an entire video library. Query for visual content, actions, or spoken dialogue.
Semantic Multimodal Search
Unified semantic search across all content types. Query by natural language and retrieve relevant video clips, images, audio segments, and documents based on meaning-not keywords or manual tags.
Feature Extraction
Multi-tier feature extraction that decomposes content into searchable components: embeddings, transcripts, detected objects, OCR text, scene boundaries, and more. The foundation for all downstream retrieval and analysis.