# Mixpeek > Multimodal AI infrastructure that makes unstructured data searchable and AI-ready. **Last Updated:** 2026-07-11 ## What is Mixpeek? Mixpeek is a multimodal data warehouse and retrieval platform for developers and AI teams. It transforms raw unstructured files (images, videos, audio, PDFs, and text) into searchable, programmable assets that AI systems can query, automate, and build upon. **Core Mission:** Infrastructure to index the world — every piece of unstructured data instantly searchable and AI-accessible. ## Quick Links - Documentation: https://mixpeek.com/docs - API Reference: https://mixpeek.com/docs/api-reference - OpenAPI Spec: https://api.mixpeek.com/docs/openapi.json - Dashboard: https://studio.mixpeek.com - Reverse Video Search: https://mixpeek.com/reverse-video-search - GitHub: https://github.com/mixpeek - Discord: https://discord.gg/mixpeek ## Machine Access (for AI agents and crawlers) - Any docs page serves raw markdown: append `.md` to its URL (e.g. https://mixpeek.com/docs/quickstart.md) - Full docs corpus in one file: https://mixpeek.com/llms-full.txt (~2 MB) - This file (content index): https://mixpeek.com/llms.txt · agent manifest: https://mixpeek.com/ai.txt - MCP: product docs for connecting agents live at https://mixpeek.com/docs/agent-integrations/mcp (.md works). Note: https://mixpeek.com/docs/mcp is a POST-only JSON-RPC MCP endpoint for docs search — a GET returns 405 by design; connect with an MCP client instead. - OpenAPI spec: https://api.mixpeek.com/docs/openapi.json ## Content Types Available | Type | Count | Path | |------|-------|------| | Blog Posts | 50+ | /blog/:slug | | Glossary Terms | 134+ | /glossary/:slug | | Recipes | 40+ | /recipes/:slug | | Technical Guides | 86 | /guides/:slug | | Curated Lists (Buyer Guides) | 49 | /curated-lists/:slug | | Feature Extractors | 20+ | /extractors/:slug | | Retrievers | 10+ | /retrievers/:slug | | Tutorials | 30+ | /tutorials/:slug | | Education Modules | 30+ | /education/module/:slug | | Comparisons | 52 | /comparisons/:slug | | Solutions | 10+ | /solutions/:slug | | Use Cases | 8+ | /use-cases/:slug | | Connectors | 10+ | /connectors/:slug | | FAQ | 52+ | /faq | ## Curated Lists (Buyer Guides & Tool Comparisons) Answer-first comparison guides for evaluating multimodal AI, vector search, and unstructured-data tooling. Best entry points for buyers comparing options. - [Best Ad Tech AI Platforms](https://mixpeek.com/curated-lists/best-ad-tech-ai-platforms) - [Best AI Content Moderation Tools](https://mixpeek.com/curated-lists/best-ai-content-moderation-tools) - [Best AI Data Warehouses](https://mixpeek.com/curated-lists/best-ai-data-warehouses) - [Best AI Digital Asset Management](https://mixpeek.com/curated-lists/best-ai-digital-asset-management) - [Best AI Ecommerce Search](https://mixpeek.com/curated-lists/best-ai-ecommerce-search) - [Best AI For Document Analysis](https://mixpeek.com/curated-lists/best-ai-for-document-analysis) - [Best AI Image Search Tools](https://mixpeek.com/curated-lists/best-ai-image-search-tools) - [Best AI Legal Document Review](https://mixpeek.com/curated-lists/best-ai-legal-document-review) - [Best AI Medical Imaging](https://mixpeek.com/curated-lists/best-ai-medical-imaging) - [Best AI Metadata Extraction Tools](https://mixpeek.com/curated-lists/best-ai-metadata-extraction-tools) - [Best AI Search Apis](https://mixpeek.com/curated-lists/best-ai-search-apis) - [Best AI Video Analysis Tools](https://mixpeek.com/curated-lists/best-ai-video-analysis-tools) - [Best AI Video Tagging Tools](https://mixpeek.com/curated-lists/best-ai-video-tagging-tools) - [Best Audio Processing Tools](https://mixpeek.com/curated-lists/best-audio-processing-tools) - [Best Brand Safety Tools](https://mixpeek.com/curated-lists/best-brand-safety-tools) - [Best Computer Vision Apis](https://mixpeek.com/curated-lists/best-computer-vision-apis) - [Best Copyright Detection Tools](https://mixpeek.com/curated-lists/best-copyright-detection-tools) - [Best Document AI Platforms](https://mixpeek.com/curated-lists/best-document-ai-platforms) - [Best Document Parsing Tools](https://mixpeek.com/curated-lists/best-document-parsing-tools) - [Best Embedding Visualization Tools](https://mixpeek.com/curated-lists/best-embedding-visualization-tools) - [Best Face Recognition Apis](https://mixpeek.com/curated-lists/best-face-recognition-apis) - [Best Feature Extraction Apis](https://mixpeek.com/curated-lists/best-feature-extraction-apis) - [Best Hybrid Search Engines](https://mixpeek.com/curated-lists/best-hybrid-search-engines) - [Best Image Recognition Apis](https://mixpeek.com/curated-lists/best-image-recognition-apis) - [Best Image Similarity Search Tools](https://mixpeek.com/curated-lists/best-image-similarity-search-tools) - [Best Image Tagging Apis](https://mixpeek.com/curated-lists/best-image-tagging-apis) - [Best Mcp Servers](https://mixpeek.com/curated-lists/best-mcp-servers) - [Best Multimodal AI Apis](https://mixpeek.com/curated-lists/best-multimodal-ai-apis) - [Best Multimodal Data Platforms](https://mixpeek.com/curated-lists/best-multimodal-data-platforms) - [Best Multimodal RAG Frameworks](https://mixpeek.com/curated-lists/best-multimodal-rag-frameworks) - [Best Multimodal Search Apis](https://mixpeek.com/curated-lists/best-multimodal-search-apis) - [Best Nsfw Detection Apis](https://mixpeek.com/curated-lists/best-nsfw-detection-apis) - [Best Object Detection Apis](https://mixpeek.com/curated-lists/best-object-detection-apis) - [Best OCR Apis](https://mixpeek.com/curated-lists/best-ocr-apis) - [Best Open Source ML Pipelines](https://mixpeek.com/curated-lists/best-open-source-ml-pipelines) - [Best PDF Extraction Tools](https://mixpeek.com/curated-lists/best-pdf-extraction-tools) - [Best RAG Frameworks](https://mixpeek.com/curated-lists/best-rag-frameworks) - [Best Reverse Image Search Apis](https://mixpeek.com/curated-lists/best-reverse-image-search-apis) - [Best Reverse Video Search Tools](https://mixpeek.com/curated-lists/best-reverse-video-search-tools) - [Best S3 Object Storage For AI](https://mixpeek.com/curated-lists/best-s3-object-storage-for-ai) - [Best Self Hosted Embedding Models](https://mixpeek.com/curated-lists/best-self-hosted-embedding-models) - [Best Semantic Search Engines](https://mixpeek.com/curated-lists/best-semantic-search-engines) - [Best Speech To Text Apis](https://mixpeek.com/curated-lists/best-speech-to-text-apis) - [Best Unstructured Data Tools](https://mixpeek.com/curated-lists/best-unstructured-data-tools) - [Best Vector Databases](https://mixpeek.com/curated-lists/best-vector-databases) - [Best Video Intelligence Apis](https://mixpeek.com/curated-lists/best-video-intelligence-apis) - [Best Video Moderation Tools](https://mixpeek.com/curated-lists/best-video-moderation-tools) - [Best Video Search Tools](https://mixpeek.com/curated-lists/best-video-search-tools) - [Best Video Transcription Tools](https://mixpeek.com/curated-lists/best-video-transcription-tools) - [Best Video Understanding Platforms](https://mixpeek.com/curated-lists/best-video-understanding-platforms) - [Best Visual Search Apis](https://mixpeek.com/curated-lists/best-visual-search-apis) ## Technical Guides (Concepts & Deep-Dives) Vendor-neutral explainers on how AI systems see, hear, and search unstructured content: embeddings, retrieval, attention, indexing, and agentic search. - [Adaptive Indexing For Agentic Search](https://mixpeek.com/guides/adaptive-indexing-for-agentic-search) - [Agent Perception Evals](https://mixpeek.com/guides/agent-perception-evals) - [Agentic Retrieval How Agents Search Differently](https://mixpeek.com/guides/agentic-retrieval-how-agents-search-differently) - [Approximate Nearest Neighbor Algorithms](https://mixpeek.com/guides/approximate-nearest-neighbor-algorithms) - [ASR Decoding Beam Search Lm Fusion](https://mixpeek.com/guides/asr-decoding-beam-search-lm-fusion) - [Audio Feature Extraction How Agents Learn To Hear](https://mixpeek.com/guides/audio-feature-extraction-how-agents-learn-to-hear) - [Audio Fingerprinting Constellation Landmark Hashing](https://mixpeek.com/guides/audio-fingerprinting-constellation-landmark-hashing) - [Audio Visual Retrieval For AI Agents](https://mixpeek.com/guides/audio-visual-retrieval-for-ai-agents) - [Bm25 Inverted Index Lexical Retrieval Internals](https://mixpeek.com/guides/bm25-inverted-index-lexical-retrieval-internals) - [Budget Aware Multi Vector Retrieval](https://mixpeek.com/guides/budget-aware-multi-vector-retrieval) - [Build Multimodal Data Warehouse](https://mixpeek.com/guides/build-multimodal-data-warehouse) - [Calibrating Similarity Scores](https://mixpeek.com/guides/calibrating-similarity-scores) - [Chunk Contextualization Late Chunking Contextual Retrieval](https://mixpeek.com/guides/chunk-contextualization-late-chunking-contextual-retrieval) - [Computer Use Agent Memory](https://mixpeek.com/guides/computer-use-agent-memory) - [Context Engineering AI Agents](https://mixpeek.com/guides/context-engineering-ai-agents) - [Contrastive Learning How Clip Siglip And Clap Work](https://mixpeek.com/guides/contrastive-learning-how-clip-siglip-and-clap-work) - [Creative Ad Analysis Agent Perception](https://mixpeek.com/guides/creative-ad-analysis-agent-perception) - [Cross Encoder Reranking](https://mixpeek.com/guides/cross-encoder-reranking) - [Deep Research Agent Over Your Own Data](https://mixpeek.com/guides/deep-research-agent-over-your-own-data) - [Diversity Aware Retrieval MMR Dpp](https://mixpeek.com/guides/diversity-aware-retrieval-mmr-dpp) - [Efficient Attention Long Context Multimodal](https://mixpeek.com/guides/efficient-attention-long-context-multimodal) - [Embedding Fine Tuning Distillation](https://mixpeek.com/guides/embedding-fine-tuning-distillation) - [Embedding Model Migration Without Reembedding](https://mixpeek.com/guides/embedding-model-migration-without-reembedding) - [Embedding Portability Versioning](https://mixpeek.com/guides/embedding-portability-versioning) - [Embedding Quantization Compression](https://mixpeek.com/guides/embedding-quantization-compression) - [Embedding Space Geometry](https://mixpeek.com/guides/embedding-space-geometry) - [Evaluating Multimodal Retrieval](https://mixpeek.com/guides/evaluating-multimodal-retrieval) - [Face Recognition Identity Clustering](https://mixpeek.com/guides/face-recognition-identity-clustering) - [Filtered Vector Search Pre Post In Place](https://mixpeek.com/guides/filtered-vector-search-pre-post-in-place) - [Forced Alignment Audio Video Agent Search](https://mixpeek.com/guides/forced-alignment-audio-video-agent-search) - [How To Check If Image Is Copyrighted](https://mixpeek.com/guides/how-to-check-if-image-is-copyrighted) - [How To Check If Picture Is Copyrighted](https://mixpeek.com/guides/how-to-check-if-picture-is-copyrighted) - [How To Check If Song Is Copyrighted](https://mixpeek.com/guides/how-to-check-if-song-is-copyrighted) - [How To Check If Video Is Copyrighted](https://mixpeek.com/guides/how-to-check-if-video-is-copyrighted) - [Hybrid Search Fusion Rrf Score Normalization](https://mixpeek.com/guides/hybrid-search-fusion-rrf-score-normalization) - [Index Freshness Incremental Updates](https://mixpeek.com/guides/index-freshness-incremental-updates) - [Instance Level Visual Matching Keypoints Geometric Verification](https://mixpeek.com/guides/instance-level-visual-matching-keypoints-geometric-verification) - [Instruction Tuned Embeddings](https://mixpeek.com/guides/instruction-tuned-embeddings) - [Late Interaction Retrieval](https://mixpeek.com/guides/late-interaction-retrieval) - [Learned Sparse Retrieval Splade Dense Hybrid](https://mixpeek.com/guides/learned-sparse-retrieval-splade-dense-hybrid) - [Long Context Video Understanding](https://mixpeek.com/guides/long-context-video-understanding) - [Mask Aware Retrieval For AI Agents](https://mixpeek.com/guides/mask-aware-retrieval-for-ai-agents) - [Matryoshka Nested Embeddings Adaptive Retrieval](https://mixpeek.com/guides/matryoshka-nested-embeddings-adaptive-retrieval) - [Mcp Tool Design For Multimodal Search](https://mixpeek.com/guides/mcp-tool-design-for-multimodal-search) - [Mcp Tools Multimodal AI Agents](https://mixpeek.com/guides/mcp-tools-multimodal-ai-agents) - [Modality Gap Cross Modal Retrieval](https://mixpeek.com/guides/modality-gap-cross-modal-retrieval) - [Monocular Depth Estimation 3d From One Image](https://mixpeek.com/guides/monocular-depth-estimation-3d-from-one-image) - [Multi Index Search Architecture](https://mixpeek.com/guides/multi-index-search-architecture) - [Multi Object Tracking Following Objects Across Frames](https://mixpeek.com/guides/multi-object-tracking-following-objects-across-frames) - [Multi Stage Retrieval How Agents Search Unstructured Data](https://mixpeek.com/guides/multi-stage-retrieval-how-agents-search-unstructured-data) - [Multimodal Chunking Strategies](https://mixpeek.com/guides/multimodal-chunking-strategies) - [Multimodal Data Warehouse Architecture](https://mixpeek.com/guides/multimodal-data-warehouse-architecture) - [Multimodal Perception For AI Agents](https://mixpeek.com/guides/multimodal-perception-for-ai-agents) - [Multimodal RAG Pipeline Architecture](https://mixpeek.com/guides/multimodal-rag-pipeline-architecture) - [Object Decomposition Layered Indexing](https://mixpeek.com/guides/object-decomposition-layered-indexing) - [OCR Document AI Internals](https://mixpeek.com/guides/ocr-document-ai-internals) - [Omnimodal Embeddings](https://mixpeek.com/guides/omnimodal-embeddings) - [Open Vocabulary Object Detection](https://mixpeek.com/guides/open-vocabulary-object-detection) - [Optical Context Compression Text As Vision Tokens](https://mixpeek.com/guides/optical-context-compression-text-as-vision-tokens) - [Payload Projection For Agentic Vector Search](https://mixpeek.com/guides/payload-projection-for-agentic-vector-search) - [Perceptual Image Hashing Near Duplicate Detection](https://mixpeek.com/guides/perceptual-image-hashing-near-duplicate-detection) - [Pre Publication Ip Clearance Guide](https://mixpeek.com/guides/pre-publication-ip-clearance-guide) - [Production Ingestion Reliability Agent Perception](https://mixpeek.com/guides/production-ingestion-reliability-agent-perception) - [Queries Without A Predefined Ontology](https://mixpeek.com/guides/queries-without-a-predefined-ontology) - [Query Transformation Pipelines For Agentic Retrieval](https://mixpeek.com/guides/query-transformation-pipelines-for-agentic-retrieval) - [Reasoning Rerankers Listwise LLM Reranking](https://mixpeek.com/guides/reasoning-rerankers-listwise-llm-reranking) - [Retrieval Control Planes For AI Agents](https://mixpeek.com/guides/retrieval-control-planes-for-ai-agents) - [Retrieval Feedback Loops Learning To Rank From Interactions](https://mixpeek.com/guides/retrieval-feedback-loops-learning-to-rank-from-interactions) - [Reverse Video Search How It Works](https://mixpeek.com/guides/reverse-video-search-how-it-works) - [Semantic Caching For Agents](https://mixpeek.com/guides/semantic-caching-for-agents) - [Speaker Diarization Who Said What](https://mixpeek.com/guides/speaker-diarization-who-said-what) - [Streaming Video Understanding Online Memory](https://mixpeek.com/guides/streaming-video-understanding-online-memory) - [Structured Extraction From Unstructured Documents](https://mixpeek.com/guides/structured-extraction-from-unstructured-documents) - [Twelvelabs Marengo Embeddings Vector Store](https://mixpeek.com/guides/twelvelabs-marengo-embeddings-vector-store) - [Vector Database Cost Comparison](https://mixpeek.com/guides/vector-database-cost-comparison) - [Vector Storage Tiering](https://mixpeek.com/guides/vector-storage-tiering) - [Video Anomaly Detection For AI Agents](https://mixpeek.com/guides/video-anomaly-detection-for-ai-agents) - [Video Frame Sampling For Embeddings](https://mixpeek.com/guides/video-frame-sampling-for-embeddings) - [Video Highlight Detection](https://mixpeek.com/guides/video-highlight-detection) - [Video Perception Layer For AI Agents](https://mixpeek.com/guides/video-perception-layer-for-ai-agents) - [Video RAG Retrieval Augmented Generation Over Video](https://mixpeek.com/guides/video-rag-retrieval-augmented-generation-over-video) - [Video Scene Segmentation](https://mixpeek.com/guides/video-scene-segmentation) - [Video Temporal Grounding](https://mixpeek.com/guides/video-temporal-grounding) - [Vision Language Model Architecture](https://mixpeek.com/guides/vision-language-model-architecture) - [Visual Document Retrieval](https://mixpeek.com/guides/visual-document-retrieval) - [What Is Multimodal Data Warehouse](https://mixpeek.com/guides/what-is-multimodal-data-warehouse) ## Head-to-Head Comparisons (Mixpeek vs Alternatives) Feature, pricing, and deployment breakdowns of Mixpeek against video AI platforms, vector databases, and search APIs. Honest about when each tool wins. - [Colpali Vs OCR](https://mixpeek.com/comparisons/colpali-vs-ocr) - [Elasticsearch Vs Pinecone](https://mixpeek.com/comparisons/elasticsearch-vs-pinecone) - [Faiss Vs Pinecone](https://mixpeek.com/comparisons/faiss-vs-pinecone) - [Google Document AI Vs Aws Textract](https://mixpeek.com/comparisons/google-document-ai-vs-aws-textract) - [Google Vision Vs Aws Rekognition](https://mixpeek.com/comparisons/google-vision-vs-aws-rekognition) - [Langchain Vs Llamaindex](https://mixpeek.com/comparisons/langchain-vs-llamaindex) - [Managed Vs Self Hosted AI](https://mixpeek.com/comparisons/managed-vs-self-hosted-ai) - [Mixpeek Vs Algolia](https://mixpeek.com/comparisons/mixpeek-vs-algolia) - [Mixpeek Vs Aws Rekognition](https://mixpeek.com/comparisons/mixpeek-vs-aws-rekognition) - [Mixpeek Vs Chroma](https://mixpeek.com/comparisons/mixpeek-vs-chroma) - [Mixpeek Vs Clarifai](https://mixpeek.com/comparisons/mixpeek-vs-clarifai) - [Mixpeek Vs Coactive](https://mixpeek.com/comparisons/mixpeek-vs-coactive) - [Mixpeek Vs Deepset](https://mixpeek.com/comparisons/mixpeek-vs-deepset) - [Mixpeek Vs Diy](https://mixpeek.com/comparisons/mixpeek-vs-diy) - [Mixpeek Vs Elasticsearch](https://mixpeek.com/comparisons/mixpeek-vs-elasticsearch) - [Mixpeek Vs Glean](https://mixpeek.com/comparisons/mixpeek-vs-glean) - [Mixpeek Vs Google Vertex AI](https://mixpeek.com/comparisons/mixpeek-vs-google-vertex-ai) - [Mixpeek Vs Haystack](https://mixpeek.com/comparisons/mixpeek-vs-haystack) - [Mixpeek Vs Hive](https://mixpeek.com/comparisons/mixpeek-vs-hive) - [Mixpeek Vs Jina](https://mixpeek.com/comparisons/mixpeek-vs-jina) - [Mixpeek Vs Lancedb](https://mixpeek.com/comparisons/mixpeek-vs-lancedb) - [Mixpeek Vs Langchain](https://mixpeek.com/comparisons/mixpeek-vs-langchain) - [Mixpeek Vs Llamaindex](https://mixpeek.com/comparisons/mixpeek-vs-llamaindex) - [Mixpeek Vs Marqo](https://mixpeek.com/comparisons/mixpeek-vs-marqo) - [Mixpeek Vs Milvus](https://mixpeek.com/comparisons/mixpeek-vs-milvus) - [Mixpeek Vs Nuclia](https://mixpeek.com/comparisons/mixpeek-vs-nuclia) - [Mixpeek Vs Pinecone](https://mixpeek.com/comparisons/mixpeek-vs-pinecone) - [Mixpeek Vs Ragie](https://mixpeek.com/comparisons/mixpeek-vs-ragie) - [Mixpeek Vs Roboflow](https://mixpeek.com/comparisons/mixpeek-vs-roboflow) - [Mixpeek Vs Turbopuffer](https://mixpeek.com/comparisons/mixpeek-vs-turbopuffer) - [Mixpeek Vs Twelvelabs](https://mixpeek.com/comparisons/mixpeek-vs-twelvelabs) - [Mixpeek Vs Typesense](https://mixpeek.com/comparisons/mixpeek-vs-typesense) - [Mixpeek Vs Unstructured](https://mixpeek.com/comparisons/mixpeek-vs-unstructured) - [Mixpeek Vs Vectara](https://mixpeek.com/comparisons/mixpeek-vs-vectara) - [Mixpeek Vs Vidrovr](https://mixpeek.com/comparisons/mixpeek-vs-vidrovr) - [Mixpeek Vs Visua](https://mixpeek.com/comparisons/mixpeek-vs-visua) - [Mixpeek Vs Weaviate](https://mixpeek.com/comparisons/mixpeek-vs-weaviate) - [Multimodal Data Warehouse Vs Data Lakehouse](https://mixpeek.com/comparisons/multimodal-data-warehouse-vs-data-lakehouse) - [Multimodal Data Warehouse Vs Multimodal Database](https://mixpeek.com/comparisons/multimodal-data-warehouse-vs-multimodal-database) - [Multimodal Data Warehouse Vs Vector Database](https://mixpeek.com/comparisons/multimodal-data-warehouse-vs-vector-database) - [Multimodal RAG Vs Text RAG](https://mixpeek.com/comparisons/multimodal-rag-vs-text-rag) - [Openai Embeddings Vs Cohere](https://mixpeek.com/comparisons/openai-embeddings-vs-cohere) - [Pinecone Vs Chroma](https://mixpeek.com/comparisons/pinecone-vs-chroma) - [Pinecone Vs Milvus](https://mixpeek.com/comparisons/pinecone-vs-milvus) - [Pinecone Vs Pgvector](https://mixpeek.com/comparisons/pinecone-vs-pgvector) - [Pinecone Vs Qdrant](https://mixpeek.com/comparisons/pinecone-vs-qdrant) - [RAG Vs Fine Tuning](https://mixpeek.com/comparisons/rag-vs-fine-tuning) - [Semantic Search Vs Keyword Search](https://mixpeek.com/comparisons/semantic-search-vs-keyword-search) - [Twelve Labs Vs Google Video Intelligence](https://mixpeek.com/comparisons/twelve-labs-vs-google-video-intelligence) - [Twelvelabs Vs Google Video Intelligence](https://mixpeek.com/comparisons/twelvelabs-vs-google-video-intelligence) - [Vector Search Vs Full Text Search](https://mixpeek.com/comparisons/vector-search-vs-full-text-search) - [Weaviate Vs Qdrant](https://mixpeek.com/comparisons/weaviate-vs-qdrant) ## Alternatives Pages (evaluating a switch) For teams evaluating alternatives to a specific vendor: what Mixpeek does differently, migration path, and honest guidance on when to stay. - [Algolia](https://mixpeek.com/alternative/algolia) - [Aws Rekognition](https://mixpeek.com/alternative/aws-rekognition) - [Chroma](https://mixpeek.com/alternative/chroma) - [Clarifai](https://mixpeek.com/alternative/clarifai) - [Coactive](https://mixpeek.com/alternative/coactive) - [Deepset](https://mixpeek.com/alternative/deepset) - [Diy](https://mixpeek.com/alternative/diy) - [Elasticsearch](https://mixpeek.com/alternative/elasticsearch) - [Glean](https://mixpeek.com/alternative/glean) - [Google Vertex AI](https://mixpeek.com/alternative/google-vertex-ai) - [Haystack](https://mixpeek.com/alternative/haystack) - [Hive](https://mixpeek.com/alternative/hive) - [Jina](https://mixpeek.com/alternative/jina) - [Lancedb](https://mixpeek.com/alternative/lancedb) - [Langchain](https://mixpeek.com/alternative/langchain) - [Llamaindex](https://mixpeek.com/alternative/llamaindex) - [Marqo](https://mixpeek.com/alternative/marqo) - [Milvus](https://mixpeek.com/alternative/milvus) - [Nuclia](https://mixpeek.com/alternative/nuclia) - [Pinecone](https://mixpeek.com/alternative/pinecone) - [Ragie](https://mixpeek.com/alternative/ragie) - [Roboflow](https://mixpeek.com/alternative/roboflow) - [Turbopuffer](https://mixpeek.com/alternative/turbopuffer) - [Twelvelabs](https://mixpeek.com/alternative/twelvelabs) - [Typesense](https://mixpeek.com/alternative/typesense) - [Unstructured](https://mixpeek.com/alternative/unstructured) - [Vectara](https://mixpeek.com/alternative/vectara) - [Vidrovr](https://mixpeek.com/alternative/vidrovr) - [Visua](https://mixpeek.com/alternative/visua) - [Weaviate](https://mixpeek.com/alternative/weaviate) ## Key Concepts - **Namespaces**: Tenant isolation boundaries — one Qdrant collection per namespace - **Buckets**: Raw file storage (maps to S3/GCS/Azure/R2/Wasabi/Tigris) - **Collections**: Processing pipelines with configured feature extractors; each collection defines what AI models run on ingested data - **Documents**: Processed assets — Qdrant point + metadata payload; `document_id` for app logic - **Retrievers**: Configurable multi-stage search pipelines; callable as tools by LLMs and agents - **Taxonomies**: Hierarchical classification systems enabling multimodal JOIN-style queries - **Clusters**: Automatic grouping and organization of semantically similar content - **Plugins**: Custom Python-based feature extractors deployed to the inference cluster - **Manifests**: Infrastructure-as-code for declarative resource management ## The 5-Stage Pipeline ### 1. Ingestion Connect and bring raw multimodal content into the platform. - Upload files directly or via presigned URL - Bucket syncs from S3, GCS, Azure Blob, Cloudflare R2, Wasabi, Tigris - Social media connectors (Instagram) - Batch processing with retry and failure tracking - Webhooks for async completion events ### 2. Extraction AI models run on ingested content to produce structured features. **Multimodal Extractors:** - Dense/sparse embeddings across text, image, video, audio **Image Extractors:** - Object detection, facial recognition, OCR, scene classification, landmark/logo detection, image segmentation **Video Extractors:** - Transcription, scene analysis, object tracking, person detection, event detection, video summarization, frame-level embeddings **Audio Extractors:** - Transcription (Whisper), speaker diarization, emotion detection, sound event detection, language identification **Text Extractors:** - Named entity recognition, keyword extraction, text classification, topic modeling, embeddings **Document Extractors:** - PDF parsing, structured text extraction, table detection **Web/Specialty Extractors:** - Web scraper, course content parser, face identity, passthrough (no-op) **Custom Plugins:** - Deploy custom Python inference code as a first-class extractor - Auto-detects dependencies, validates manifest features at deploy time - Stable SDK: `from shared.plugins import open_asset, download_asset` - `POST /{plugin_id}/realtime/test` for live debugging ### 3. Enrichment Apply structure and context to processed data. - **Taxonomies**: Hierarchical classification with versioning and analytics - **Clusters**: Automatic semantic clustering with scheduled triggers - **Retriever-based enrichment**: Use a retriever output to enrich documents - **LLM enrichment**: GPT/Claude-powered field generation on retrieval results ### 4. Indexing Vectors and metadata stored in tiered storage. - Hot tier: Qdrant (sub-100ms vector search) - Cold tier: S3 Vectors (cost-optimized) - Archived: Metadata only - Storage tiering lifecycle: active → cold → archived - Document lineage tracking (source object → derived documents → decomposition tree) ### 5. Retrieval Multi-stage configurable retriever pipelines. **Filter Stages:** - `feature_search` — vector/semantic search against one or more feature embeddings - `attribute_filter` — metadata/payload filtering (MongoDB-style operators) - `llm_filter` — LLM-based semantic filter on candidate results - `agent_search` — autonomous agent-driven search - `query_expand` — expand query into multiple sub-queries **Sort Stages:** - `sort_relevance` — score-based ordering - `sort_attribute` — field-based ordering - `mmr` — Maximal Marginal Relevance for diversity - `rerank` — cross-encoder reranking - `score_normalize` — normalize scores across stages **Reduce Stages:** - `aggregate` — group and reduce results - `sample` — random sampling - `summarize` — LLM summarization of result set - `limit` — top-K cutoff - `deduplicate` — remove near-duplicate results **Group Stages:** - `group_by` — bucket results by field value - `cluster` — semantic clustering of result set **Apply Stages:** - `json_transform` — reshape/project fields - `rag_prepare` — format results for RAG context injection - `external_web_search` — augment with live web results - `api_call` — call an external HTTP endpoint - `sql_lookup` — join with a SQL data source - `cross_compare` — LLM-powered comparison across results - `web_scrape` — fetch and extract content from URLs in results - `unwind` — flatten array fields - `code_execution` — run sandboxed Python on results **Enrich Stages:** - `llm_enrich` — add LLM-generated fields to each result - `taxonomy_enrich` — apply taxonomy classification - `document_enrich` — join with related documents - `agentic_enrich` — autonomous enrichment via agent ## Retriever Features - **Executions**: Full execution history with explain plan - **Evaluations**: Dataset-based retrieval quality measurement (NDCG, MRR, Recall) - **Interactions**: Track clicks, views, conversions for relevance learning - **Fusion strategies**: Weighted, RRF, learned fusion across stages - **Benchmarks**: Compare retriever configurations head-to-head - **Adhoc execution**: One-off retriever runs without saving configuration - **Published retrievers**: Share retrievers publicly (callable without auth) - **Marketplace**: Publish and subscribe to community retrievers and plugins - **Auto-optimization**: Automatic parameter tuning based on interaction data ## Relevance & Personalization - Interaction tracking (clicks, views, ratings, purchases) - Learned fusion: ML-trained weights from interaction signals - Evaluation datasets with ground truth for offline testing - Analytics dashboard for retrieval quality over time ## Organization & Multi-Tenancy - Organizations with users, roles, and RBAC - API key management per user - Connections to external data sources - Usage metering and billing - Audit logs - Secrets management - Webhooks (async event delivery) - Alerts (threshold-based, event-based) ## Templates & Automation - **Templates**: Reusable configurations for namespaces, retrievers, collections, clusters, buckets, taxonomies - **Scaffolds**: Single-call namespace provisioning from a template - **Quickstart**: One API call to provision a complete working environment - **Manifest (IaC)**: Declarative YAML/JSON definitions — apply, validate, export, diff - **Cluster triggers**: Scheduled automatic re-clustering - **Bucket syncs**: Scheduled or event-driven sync from external storage - **Tasks API**: Track and manage all async operations ## Custom Models - Upload and deploy custom ML models to the inference cluster - Namespace-scoped and org-scoped model registries - Deploy to Ray object store for low-latency inference ## Developer Tools - **Python SDK**: `pip install mixpeek` - **JavaScript SDK**: `npm install mixpeek-sdk` - **CLI**: `pip install mixpeek-cli` - **MCP Server**: Expose retrievers as MCP tools for Claude/agents - **OpenAPI**: Full Swagger spec at https://api.mixpeek.com/docs/openapi.json - **Webhooks**: Async delivery for ingestion, batch, cluster, taxonomy events - **Resource Search**: Cross-resource search across all platform entities ## API Authentication All API requests require Bearer token authentication: ``` Authorization: Bearer YOUR_API_KEY ``` Namespace context is passed via header: ``` X-Namespace: ns_your_namespace_id ``` ## Example: Ingest and Search ```python from mixpeek import Mixpeek client = Mixpeek(api_key="your-api-key") # Upload a file to a bucket upload = client.buckets.uploads.create( bucket_id="my-bucket", file_path="video.mp4" ) # Collection processes it automatically via configured extractors # Poll batch status or listen for webhook # Execute a retriever pipeline results = client.retrievers.execute( retriever_id="my-retriever", inputs={"query": "person walking in the park"} ) ``` ## Supported Storage Integrations - AWS S3 - Google Cloud Storage (GCS) - Azure Blob Storage - Cloudflare R2 - Wasabi - Tigris - Direct file upload - Instagram (social media connector) ## Pricing Usage-based (v2, live since 2026-07). Three questions set your bill: what kind of files (video, image, audio, documents, text, web), how much of them (billed in natural units), and what you want to search by (each enabled feature is metered on the same unit as its file type). ### Plans | Plan | Price | Includes | |---|---|---| | Build | $25/mo | $10/mo usage pool (about 200 video minutes or 6,667 images) | | Scale | $250/mo | $100/mo usage pool (about 2,000 video minutes or 66,667 images) | | Enterprise | Custom | Single-tenant deployment available (dedicated, isolated data plane) | | MVS standalone | from $25/mo | Bring-your-own-vectors vector store on object storage | ### Usage Rates | Unit | Rate | |---|---| | Video | $0.05 per minute | | Audio | $0.01 per minute | | Images | $1.50 per 1,000 | | Document pages | $1.50 per 1,000 | | Text | $2 per 1M tokens | | Web pages | $20 per 1,000 | | Feature storage | $0.33 per GB-month | | Search queries | $2 per 1M (overage) | Full-resolution processing is 2x the base rate. Additional search features (faces, on-screen text, transcripts, objects) are add-on rates on the same unit; clustering is included at no extra charge. Get an authoritative preflight quote for any batch: POST /v1/organizations/billing/estimate. Live rate card and calculator: https://mixpeek.com/pricing ## Deployment - Fully managed cloud (recommended) - Enterprise single-tenant: dedicated, isolated data plane (database, compute, vector shard, bucket) in your chosen cloud and region - Bring your own object storage (S3/GCS): Mixpeek indexes content in place ## Technical Performance - Sub-100ms inference with model caching and GPU acceleration - Concurrent processing of thousands of documents - Auto-scaling Ray-based inference cluster - Support for large files (multi-hour video) ## Use Cases by Industry - **Advertising & Media**: Contextual targeting, brand safety, creative analysis - **Entertainment**: Content discovery, video library search, recommendations - **E-commerce**: Visual product search, automated tagging, inventory tracking - **Security**: AI threat detection, anomaly detection, incident analysis - **Healthcare**: Medical imaging analysis, patient data integration - **EdTech**: Smart content management, adaptive learning, learning analytics - **Manufacturing**: Safety compliance, defect detection, predictive maintenance - **Legal**: Multimodal eDiscovery, compliance monitoring - **Dataset Engineering**: Intelligent curation, annotation workflows, AI training data ## Why Mixpeek? ### vs. Vector Databases (Pinecone, Weaviate, Qdrant) - End-to-end solution including feature extraction — not just storage - Multi-modal native, not just text - Managed ML inference included ### vs. ML Platforms (Vertex AI, SageMaker) - Specialized for retrieval use cases - Higher-level API, less infrastructure management - Integrated pipeline in one platform ### vs. Search Platforms (Elasticsearch, Algolia) - AI-native semantic/vector search - Native multimodal support - Feature extraction built-in ### vs. Building In-House - 12-18 months engineering time saved - Production-ready with monitoring, error handling, auto-scaling - Cost optimization with managed GPU resources ## Getting Started 1. **Sign up**: https://studio.mixpeek.com 2. **Get API key**: From dashboard settings 3. **Create namespace**: Isolated tenant environment 4. **Create bucket + collection**: Configure feature extractors 5. **Upload content**: Files processed automatically 6. **Create retriever**: Configure search pipeline stages 7. **Execute retriever**: Query with natural language or multimodal input ## Support & Resources - Documentation: https://mixpeek.com/docs - API Reference: https://mixpeek.com/docs/api-reference - Recipes & Examples: https://mixpeek.com/recipes - Blog: https://mixpeek.com/blog - GitHub: https://github.com/mixpeek - Discord: https://discord.gg/mixpeek - Email: support@mixpeek.com ## Company - **Founded**: 2023 - **Headquarters**: New York, NY - **Founder & CEO**: Ethan Steininger (former MongoDB Search Team Lead) - **Investors**: Work-Bench, Essence VC, Humans of the Internet --- *This document follows the llms.txt standard for providing structured information to AI systems and crawlers about Mixpeek's capabilities.*