AI-Powered Catalog Search
Replace keyword-based catalog search with AI-powered semantic and visual search. Understands natural language queries like "lightweight summer dress under $50" and combines text understanding with visual similarity for comprehensive product discovery.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="catalog")# 1. A bucket whose objects carry a product photo and its description, and one collection per extractor over itbucket = client.buckets.create(bucket_name="products",bucket_schema={"properties": {"photo": {"type": "image",},"description": {"type": "text",},},},)image_collection = client.collections.create(collection_name="product_photos",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "multimodal_extractor","version": "v1",},)text_collection = client.collections.create(collection_name="product_text",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "text_extractor","version": "v1",},)# 2. Upload and processclient.buckets.upload(bucket["bucket_id"],blobs=[{"property": "photo", "type": "image", "data": "s3://your-bucket/products/linen-dress.jpg"}, {"property": "description", "type": "text", "data": "s3://your-bucket/products/linen-dress.txt"}],)client.collections.trigger(image_collection["collection_id"])client.collections.trigger(text_collection["collection_id"])# 3. The query searched against photos and descriptions, fused with RRF, then a rerank on description textretriever = client.retrievers.create(retriever_name="catalog-search",collection_identifiers=["product_photos", "product_text"],input_schema={"query": {"type": "text","required": True,},},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://multimodal_extractor@v1/vertex_multimodal_embedding","query": {"input_mode": "text","value": "{{INPUT.query}}",},"top_k": 100,},{"feature_uri": "mixpeek://text_extractor@v1/multilingual_e5_large_instruct_v1","query": {"input_mode": "text","value": "{{INPUT.query}}",},"top_k": 100,},],"fusion": "rrf","final_top_k": 100,},},{"stage_name": "rerank","stage_id": "rerank","parameters": {"inference_name": "BAAI__bge_reranker_v2_m3","query": "{{INPUT.query}}","document_field": "text","top_k": 20,},},],)# 4. Searchresults = client.retrievers.execute(retriever["retriever_id"],inputs={"query": "lightweight summer dress under $50",},)for doc in results["documents"]:print(doc["document_id"], doc["score"])
Feature Extractors
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Text Embedding
Extract semantic embeddings from documents, transcripts and text content
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
rerank
Rerank documents using cross-encoder models for accurate relevance
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
Text Embedding
Extract semantic embeddings from documents, transcripts and text content
Visual Product Search
Enable camera-based product discovery for e-commerce. Customers snap a photo of a product they like, and the pipeline returns visually similar items from your catalog along with pricing, availability, and product metadata.
Multimodal Extractor
Unified embeddings for video, audio, image, and text: scene/silence chunking, Whisper transcription, thumbnails, and Gemini vision.
Reverse Image Search
Build a reverse image search engine that accepts an image as input and returns visually similar results from your indexed collection. Uses multimodal embeddings to match visual features like color, shape, composition, and semantic content.
E-commerce Catalog Enrichment
Automatically enrich product listings with visual attributes, descriptions, and tags extracted from product images. Fill in missing catalog data at scale.
Face Search Pipeline
Detect, align, and embed faces across your image and video library, then search by face identity. Uses SCRFD for detection and ArcFace for 512-dimensional identity embeddings, enabling large-scale face recognition and matching.