Document Intelligence Search
Extract and search through PDFs, presentations, and documents. Combines OCR, layout analysis, and semantic search for comprehensive document retrieval.
from mixpeek import Mixpeekclient = Mixpeek(api_key="YOUR_API_KEY", namespace="doc-search")# 1. A bucket for the PDFs, and a collection that splits pages into layout blocks with OCR, confidence and an E5 embedding eachbucket = client.buckets.create(bucket_name="contracts",bucket_schema={"properties": {"contract": {"type": "pdf",},},},)collection = client.collections.create(collection_name="contracts",source={"type": "bucket", "bucket_ids": [bucket["bucket_id"]]},feature_extractor={"feature_extractor_name": "document_graph_extractor","version": "v1",},)# 2. Upload and processclient.buckets.upload(bucket["bucket_id"],blobs=[{"property": "contract", "type": "pdf", "data": "s3://your-bucket/contracts/msa-acme.pdf"}],)client.collections.trigger(collection["collection_id"])# 3. Dense and BM25 search over block text fused with RRF, then a cross-encoder rerank (BM25 reads the namespace text payload indexes)retriever = client.retrievers.create(retriever_name="contract-search",collection_identifiers=["contracts"],input_schema={"query": {"type": "text","required": True,},},stages=[{"stage_name": "search","stage_id": "feature_search","parameters": {"searches": [{"feature_uri": "mixpeek://document_graph_extractor@v1/intfloat__multilingual_e5_large_instruct","query": {"input_mode": "text","value": "{{INPUT.query}}",},"top_k": 50,},{"feature_uri": "mixpeek://document_graph_extractor@v1/intfloat__multilingual_e5_large_instruct","query": {"input_mode": "text","value": "{{INPUT.query}}",},"top_k": 50,"lexical": True,},],"fusion": "rrf","final_top_k": 50,},},{"stage_name": "rerank","stage_id": "rerank","parameters": {"inference_name": "BAAI__bge_reranker_v2_m3","query": "{{INPUT.query}}","document_field": "text_raw","top_k": 10,},},],)# 4. Searchresults = client.retrievers.execute(retriever["retriever_id"],inputs={"query": "indemnification clause with liability cap",},)for doc in results["documents"]:print(doc["page_number"], doc["text_raw"][:80], doc["score"])
Feature Extractors
Document Graph Extractor
Decompose PDFs into spatial blocks (paragraphs, tables, forms, headers) with layout classification and E5 text embeddings.
Retriever Stages
feature search
Search and filter documents by vector similarity using feature embeddings
rerank
Rerank documents using cross-encoder models for accurate relevance
Use Cases Using This Recipe
Insurance Claims Document Processing
Extract structured data from claims documents, photos, and correspondence automatically
70% reduction in manual document handling
Adjuster data entry time
Insurance carriers, claims adjusters, and third-party administrators processing 1,000+ claims monthly across property, casualty, auto, and health lines
Semantic Search for Knowledge Bases
Find answers by meaning, not keywords, across your entire knowledge repository
85% of queries answered on first search vs. 40% baseline
First-search success rate
Knowledge management teams, internal documentation owners, customer support organizations, and EdTech platforms maintaining 10K+ articles, documents, and multimedia resources
Enterprise RAG Search
Ask questions across all your enterprise data and get sourced, verifiable answers
80% faster from question to answer
Information retrieval time
Financial services firms, consulting organizations, legal teams, and enterprise knowledge workers who need to synthesize information across thousands of internal documents, reports, and presentations
Clinical NLP at Scale
Extract structured intelligence from clinical notes, pathology reports, and medical records
94% F1 on medical NER benchmarks
Entity extraction accuracy
Healthcare IT teams, clinical informatics departments, and health systems processing thousands of clinical documents daily
Related Recipes & Resources
Explore these related resources to deepen your understanding and discover more powerful features
BYO Embeddings Vector Search
Bring pre-computed embeddings from any provider (OpenAI, Cohere, Together, etc.) and upsert them directly into MVS for instant vector search. No feature extractors, no pipelines -- just embeddings in, results out.
Document Graph Extractor
Decompose PDFs into spatial blocks (paragraphs, tables, forms, headers) with layout classification and E5 text embeddings.
Multimodal Hybrid Search Pipeline
Combine vector search with keyword search (BM25) across text, images, and video for the most comprehensive multimodal retrieval system.
Clinical Documentation Structuring
Production-grade pipeline for ingesting clinical documents, scanned charts, EHR exports, wound photos, and therapy notes, and structuring them into coded fields aligned with MDS 3.0, PDPM, and CMS audit requirements. Combines OCR, clinical NER, taxonomy classification, and hybrid retrieval to turn unstructured bedside documentation into queryable, auditable data.
Dense Search Over Your Own Embeddings, and What Hybrid Needs
Upsert documents you embedded elsewhere into an MVS namespace and search them by raw vector through the features search endpoint. Hybrid BM25 plus dense is not part of a plain BYO upsert: the documents carry dense vectors only and no text index is created. If you want a lexical leg later, declare a TEXT payload index on the field when you create the namespace; this recipe shows that declaration and the dense search that works today.
Multimodal Search with MVS
Build multimodal search by embedding different content types (text, images, video frames) with your own models and searching across them in a single MVS namespace. Use CLIP or any multimodal embedding model for cross-modal retrieval.