Semantic Drift Detection
Monitor distribution shifts between baseline and current data using cluster comparison. Detect when new content diverges from training data or when content mix changes unexpectedly.
"Detect distribution drift in training data between Q1 baseline and current dataset"
Why This Matters
Data drift is silent model degradation. By comparing cluster distributions over time, you catch drift before it impacts production systems.
import timeimport requestsfrom mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="drift-monitor")API = "https://api.mixpeek.com/v1"HEADERS = {"Authorization": "Bearer YOUR_API_KEY", "X-Namespace": "drift-monitor"}# 1. The training images, embeddedbucket = client.buckets.create(bucket_name="training-data",bucket_schema={"properties": {"image": {"type": "image"}}},)collection = client.collections.create(collection_name="training-data",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1",},)client.buckets.upload(bucket["bucket_id"],blobs=[{"property": "image", "type": "image", "data": "s3://your-bucket/training/batch-01/img-0001.jpg"}],)client.collections.trigger(collection["collection_id"])# 2. One cluster definition, run whenever you want a snapshot of the embedding space.# The SDK has no clusters resource, so this part is REST.cluster = requests.post(API + "/clusters", headers=HEADERS, json={"cluster_name": "training-drift","collection_ids": [collection["collection_id"]],"cluster_type": "vector","vector_config": {"feature_uris": ["mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"],"clustering_method": "hdbscan","algorithm_params": {"min_cluster_size": 20},},}).json()def snapshot():task = requests.post(API + "/clusters/" + cluster["cluster_id"] + "/execute", headers=HEADERS, json={}).json()while client.tasks.get(task["task_id"])["status"] not in ("COMPLETED", "COMPLETED_WITH_ERRORS", "FAILED"):time.sleep(15)return requests.get(API + "/clusters/" + cluster["cluster_id"] + "/executions", headers=HEADERS).json()baseline = snapshot()# 3. After new data lands and is processed, snapshot again and compare the runsclient.buckets.upload(bucket["bucket_id"],blobs=[{"property": "image", "type": "image", "data": "s3://your-bucket/training/batch-02/img-0001.jpg"}],)client.collections.trigger(collection["collection_id"])current = snapshot()for name in ("noise_ratio", "silhouette_score", "mean_cosine_to_centroid"):before = (baseline.get("metrics") or {}).get(name)after = (current.get("metrics") or {}).get(name)print(name, before, "->", after)print("clusters:", baseline["num_clusters"], "->", current["num_clusters"])# Every run stays in the historyhistory = requests.post(API + "/clusters/" + cluster["cluster_id"] + "/executions/list", headers=HEADERS, json={}).json()for run in history["results"]:print(run["run_id"], run["num_clusters"], (run.get("metrics") or {}).get("noise_ratio"))
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
Documentation
Use Cases Using This Recipe
Creative Lineage & Storyboard Intelligence
Track creative evolution from concept to final cut
85% concept retention
Brief-to-final alignment
Creative directors, brand managers, and production teams managing multi-version creative workflows
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Dataset Versioning
Treat versioned object storage as your dataset's source of truth. Capture complete snapshots-raw assets, embeddings, and cluster assignments-for deterministic reconstruction at any point in time.
Clustering & Theme Discovery
Unsupervised clustering that groups content into semantic themes using HDBSCAN. Surfaces hidden patterns, content variants, and outliers without requiring predefined labels.
Hierarchical Classification
Assign content to multi-level category hierarchies using embedding-based classification. Define your taxonomy once, then classify new content automatically with confidence scores.
Brand Safety & Ad Verification Pipeline
GARM-compliant brand safety pipeline for ad networks. Analyze video and image creatives for brand safety violations before serving.
Automated Video Tagging
Automatically generate descriptive tags for video content using scene analysis, object detection, and taxonomy classification. Each video receives structured labels for scenes, objects, actions, and custom business categories without manual annotation.