Clustering & Theme Discovery
Unsupervised clustering that groups content into semantic themes using HDBSCAN. Surfaces hidden patterns, content variants, and outliers without requiring predefined labels.
"Discover hidden themes in unlabeled user-generated content and identify outliers"
Why This Matters
You can't search for what you don't know exists. Clustering reveals the natural structure in your content-themes, duplicates, and anomalies-before you even ask.
import timeimport requestsfrom mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="content-themes")API = "https://api.mixpeek.com/v1"HEADERS = {"Authorization": "Bearer YOUR_API_KEY", "X-Namespace": "content-themes"}# 1. A collection that embeds every imagebucket = client.buckets.create(bucket_name="content",bucket_schema={"properties": {"image": {"type": "image"}}},)collection = client.collections.create(collection_name="content",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1",},)client.buckets.upload(bucket["bucket_id"],blobs=[{"property": "image", "type": "image", "data": "s3://your-bucket/content/post-0001.jpg"}],)client.collections.trigger(collection["collection_id"])# 2. HDBSCAN finds dense groups without a preset count, and an LLM names each one.# The SDK has no clusters resource, so this part is REST.cluster = requests.post(API + "/clusters", headers=HEADERS, json={"cluster_name": "content-themes","collection_ids": [collection["collection_id"]],"cluster_type": "vector","vector_config": {"feature_uris": ["mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding"],"clustering_method": "hdbscan","algorithm_params": {"min_cluster_size": 15},},"llm_labeling": {"enabled": True, "provider": "openai", "model_name": "gpt-4o-mini-2024-07-18"},}).json()task = requests.post(API + "/clusters/" + cluster["cluster_id"] + "/execute", headers=HEADERS, json={}).json()while client.tasks.get(task["task_id"])["status"] not in ("COMPLETED", "COMPLETED_WITH_ERRORS", "FAILED"):time.sleep(15)# 3. Each group carries its label, size and keywordsgroups = requests.get(API + "/clusters/" + cluster["cluster_id"] + "/groups", headers=HEADERS).json()for group in groups["groups"]:print(group["label"], group["member_count"], ", ".join(group.get("keywords") or []))
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
Documentation
Use Cases Using This Recipe
Multimodal Lead Intelligence
Enrich leads with visual and behavioral signals from their content
+30% improvement
Lead scoring accuracy
B2B sales teams, demand gen marketers, and ABM platforms enriching 10K+ leads monthly
Talent Intelligence & Casting
Match talent to roles using multimodal portfolio analysis
75% reduction
Casting search time
Casting directors, talent agencies, and production companies managing 10K+ talent profiles
Social Media Content Intelligence
Analyze and optimize social content performance with multimodal AI
+35% average improvement
Content engagement rate
Social media managers, content strategists, and brand teams publishing 100+ posts monthly across platforms
Fashion Visual Product Discovery
Search for fashion by style, not just by name or brand
3x more products viewed per session
Product discovery engagement
Fashion e-commerce platforms, apparel retailers, and personal styling services managing catalogs of 100K+ products where visual style drives purchase decisions
AI-Powered Stock Media Search
Find the perfect stock asset by describing what you envision, not what keywords to try
+45% more purchases per search session
Search-to-license conversion rate
Stock media platforms, content licensing marketplaces, and enterprise media libraries serving creative professionals who need to find specific visual and audio assets quickly
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Dataset Versioning
Treat versioned object storage as your dataset's source of truth. Capture complete snapshots-raw assets, embeddings, and cluster assignments-for deterministic reconstruction at any point in time.
Semantic Multimodal Search
Unified semantic search across all content types. Query by natural language and retrieve relevant video clips, images, audio segments, and documents based on meaning-not keywords or manual tags.
Feature Extraction
Multi-tier feature extraction that decomposes content into searchable components: embeddings, transcripts, detected objects, OCR text, scene boundaries, and more. The foundation for all downstream retrieval and analysis.
Anomaly Detection
Identify outliers and anomalous content using embedding distance from cluster centroids. Flag quality issues, novel content, or items that don't match expected patterns.
Semantic Drift Detection
Monitor distribution shifts between baseline and current data using cluster comparison. Detect when new content diverges from training data or when content mix changes unexpectedly.