Video Content Analytics Pipeline
Analyze video content at scale to extract insights: scene composition, speaker time, topic distribution, and sentiment across your video library.
from collections import Counterfrom mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="video-analytics")# 1. A bucket for campaign videos, and a collection that cuts at scene changes,# transcribes, and has Gemini return topic and sentiment per scenebucket = client.buckets.create(bucket_name="campaign-videos",bucket_schema={"properties": {"video": {"type": "video"}}},)collection = client.collections.create(collection_name="marketing-videos",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"split_method": "scene","run_transcription": True,"run_video_description": True,"response_shape": {"type": "object","properties": {"topic": {"type": "string"},"sentiment": {"type": "string", "enum": ["positive", "neutral", "negative"]},},},},},)# 2. Upload and processclient.buckets.upload(bucket["bucket_id"],blobs=[{"property": "video", "type": "video", "data": "s3://your-bucket/campaign-videos/spot-01.mp4"}],)client.collections.trigger(collection["collection_id"])# 3. Aggregate across scenesdocs = client.documents.list(collection["collection_id"], page_size=500)sentiment = Counter((doc.get("json_output") or {}).get("sentiment") for doc in docs["results"])seconds_by_topic = Counter()for doc in docs["results"]:topic = (doc.get("json_output") or {}).get("topic", "unknown")seconds_by_topic[topic] += doc["end_time"] - doc["start_time"]print(sentiment.most_common(), seconds_by_topic.most_common(5))
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
Related Blog Posts
Use Cases Using This Recipe
AI Video Surveillance Analytics
Transform passive camera feeds into actionable security intelligence
85% of events caught live vs. 5% manual baseline
Real-time incident detection rate
Security operations centers, facility managers, and enterprise security teams monitoring 50+ camera feeds across multiple locations
Video Analytics for Sports Broadcasting
Unlock play-by-play intelligence from broadcast footage at scale
Seconds instead of hours per clip
Moment discovery time
Sports broadcasters, league media teams, sports analytics companies, and OTT platforms managing multi-season video archives across multiple sports
Automated Video Tagging for Streaming
Auto-generate rich metadata for every scene, shot, and moment in your catalog
10x more tags than manual editorial process
Metadata tags per title
Streaming platforms, content distributors, and VOD services managing catalogs of 10K+ titles that need rich metadata for discovery and recommendation
Video Compliance Monitoring
Automatically verify regulatory and policy compliance across video content at scale
10x faster than manual review
Compliance review speed
Compliance teams, broadcast standards departments, advertising regulators, and enterprise communications teams reviewing video content for regulatory adherence
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Hierarchical Classification
Assign content to multi-level category hierarchies using embedding-based classification. Define your taxonomy once, then classify new content automatically with confidence scores.
Video Transcription & Indexing Pipeline
Automatically transcribe video content with speaker identification, timestamps, and full-text indexing for downstream search and analytics.
Video RAG Pipeline
Retrieval-augmented generation specifically designed for video content. Decomposes videos into scenes and transcripts, retrieves relevant segments for a given question, and passes them as context to an LLM with precise timestamp citations.
Searchable Video Library
Turn an unstructured video archive into a fully searchable library. Each video is decomposed into scenes with transcriptions, visual embeddings, and metadata. Users search by natural language and jump directly to the relevant moment in any video.
Semantic Join
Bridge extracted content features with business reference data. Join video clips to product catalogs, detected faces to employee directories, or documents to compliance frameworks-all via embedding similarity.