Automatic Image Tagging Pipeline
Automatically tag images with objects, scenes, colors, and custom labels using computer vision models. Perfect for DAM systems and media libraries.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="image-tags")# 1. A bucket for the photos, and a collection whose Gemini description step# returns tags as structured JSON. There is no separate object detector or# color extractor; one structured answer per image carries all three.bucket = client.buckets.create(bucket_name="product-photos",bucket_schema={"properties": {"photo": {"type": "image"}}},)collection = client.collections.create(collection_name="product-photos",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1","parameters": {"run_video_description": True,"description_prompt": "Tag this photo for a media library.","response_shape": {"type": "object","properties": {"tags": {"type": "array", "items": {"type": "string"}},"objects": {"type": "array", "items": {"type": "string"}},"dominant_colors": {"type": "array", "items": {"type": "string"}},},},},},)# 2. Upload and processclient.buckets.upload(bucket["bucket_id"],blobs=[{"property": "photo", "type": "image", "data": "s3://your-bucket/photos/IMG_0001.jpg"}],)client.collections.trigger(collection["collection_id"])# 3. Read the tags backdocs = client.documents.list(collection["collection_id"], page_size=50)for doc in docs["results"]:print(doc["document_id"], doc.get("json_output"))
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Retriever Stages
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Image Embedding
Generate visual embeddings for similarity search and clustering
Object Detection
Identify and locate objects within images with bounding boxes
Image Captioning
Generate descriptive captions for images automatically
Facial Recognition
Detect and identify faces in images with high accuracy
Image Segmentation
Partition images into multiple segments for detailed analysis