Multimodal Content Moderation
Automated content moderation pipeline that analyzes text, images, and video for policy violations. Uses hierarchical taxonomy classification to label content as safe, sensitive, or prohibited across multiple categories simultaneously.
import requestsfrom mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="moderation")API = "https://api.mixpeek.com/v1"HEADERS = {"Authorization": "Bearer YOUR_API_KEY", "X-Namespace": "moderation"}# 1. Policy exemplars: images labeled safe, sensitive or prohibitedexample_bucket = client.buckets.create(bucket_name="policy-examples",bucket_schema={"properties": {"image": {"type": "image"}}},)examples = client.collections.create(collection_name="policy-examples",source={"type": "bucket", "bucket_ids": [example_bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1",},)client.buckets.upload(example_bucket["bucket_id"],blobs=[{"property": "image", "type": "image", "data": "s3://your-bucket/policy/prohibited-weapons-01.jpg"}],)client.collections.trigger(examples["collection_id"])# 2. The retriever the taxonomy uses to find the closest exemplarmatcher = client.retrievers.create(retriever_name="policy-matcher",collection_identifiers=["policy-examples"],input_schema={"query": {"type": "text", "required": True}},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding","query": {"input_mode": "text", "value": "{{INPUT.query}}"},"top_k": 3,},],"final_top_k": 3,},},],)# A flat taxonomy matches each document, through a retriever, against a# collection of labeled examples. The SDK has no taxonomies resource, so this is REST.taxonomy = requests.post(API + "/taxonomies", headers=HEADERS, json={"taxonomy_name": "content_moderation","config": {"taxonomy_type": "flat","retriever_id": matcher["retriever_id"],"input_mappings": [{"input_key": "query", "source_type": "payload", "path": "description"}],"source_collection": {"collection_id": examples["collection_id"]},},}).json()# 3. User content, labeled at query time, with an LLM check for the borderline casesbucket = client.buckets.create(bucket_name="user-content",bucket_schema={"properties": {"upload": {"type": "image"}}},)uploads = client.collections.create(collection_name="user-content",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"run_video_description": True,},},)moderation = client.retrievers.create(retriever_name="moderation-review",collection_identifiers=["user-content"],input_schema={"image": {"type": "image", "required": True}},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding","query": {"input_mode": "content", "value": "{{INPUT.image}}"},"top_k": 20,},],"final_top_k": 20,},},{"stage_name": "label", "stage_id": "taxonomy_enrich", "parameters": {"taxonomy_id": taxonomy["taxonomy_id"], "top_k": 1}},{"stage_name": "borderline", "stage_id": "llm_filter", "parameters": {"criteria": "Keep content that shows weapons, nudity or hate symbols.", "provider": "google", "model_name": "gemini-2.5-flash-lite"}},],)results = client.retrievers.execute(moderation["retriever_id"], inputs={"image": "https://example.com/user-upload.jpg"})for doc in results["documents"]:print(doc["document_id"], doc["score"])
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
taxonomy enrich
Classify documents against taxonomy nodes via vector similarity
llm filter
Filter documents using LLM-based semantic evaluation
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Document Classification Pipeline
Classify documents into custom business categories using layout-aware extraction and taxonomy enrichment. Handles invoices, contracts, reports, forms, and correspondence by analyzing both textual content and visual document structure.
Multimodal Hybrid Search Pipeline
Combine vector search with keyword search (BM25) across text, images, and video for the most comprehensive multimodal retrieval system.
Multimodal RAG Pipeline
Build a retrieval-augmented generation system that works with text, images, and video. Feed relevant multimodal context to LLMs for grounded responses.
Taxonomy Enrichment Pipeline
Automatically classify and tag content using custom taxonomies. Map your content to IAB categories, custom hierarchies, or industry-specific classifications.
Content Clustering Pipeline
Automatically group similar content together using embedding-based clustering. Discover themes, identify duplicates, and organize large content libraries.