Browse the extractor catalog on GitHub
Runnable reference for every built-in Mixpeek extractor — inputs, parameters, output fields, embedding models, and copy-paste examples. Auto-generated from the live registry, so it always matches production.
Availability
What You Can Build
Custom extractors plug into two places in the warehouse:1. Feature Extractors (Decomposition)
Attach custom logic to a collection so every ingested object flows through your pipeline during decomposition. Your extractor is your vocabulary — unlike built-in pipelines (which are selected via features), custom extractors keep their explicit names and are selected as a custom feature:feature_extractor config (shown later on this page) also works — use it when you need custom input_mappings or parameters. Use custom extractors to:
- Embed domain-specific content with your own model (fine-tuned CLIP, proprietary audio encoder, etc.)
- Extract structured attributes via a VLM you manage (brand compliance, regulated content classification)
- Transcribe, OCR, or segment media with a custom pre/post-processing chain
- Produce multiple named vector indexes from a single pass
mixpeek://my_extractor@1.0.0/my_embedding), so retrievers, taxonomies, and clusters can reference them.
Pricing: a
custom:<name> feature derives its per-unit rate from the compute profile your extractor declares — the same machinery that prices native features. Quote it before running via POST /v1/organizations/billing/estimate; see Billing & Pricing.2. Retriever Operations (Query Time)
An extractor’srealtime.py exposes a Ray Serve HTTP endpoint that retriever stages can call during execution. Use this to power:
feature_search— embed queries at search time with the same model you used during ingestion, so the query vector lives in the same space as the indexed vectors- Inference operations — on-the-fly classification, scoring, or re-ranking against your model
- LLM calls — wrap a hosted or private LLM behind a stable contract, with platform-managed secrets and cost tracking
- Classifiers — apply your own classifier to candidate results mid-pipeline
Delivery Formats
Ship your extractor in either of two formats:
Both formats expose the same runtime APIs (batch
__call__, real-time run_inference, platform LLM/secret accessors).
Extractor Structure
Every extractor has the same layout:manifest.pydeclares what your extractor accepts, produces, and which vector indexes to createpipeline.pywires your processor into the Ray Data batch pipelinerealtime.pyexposes a Ray Serve endpoint for query-time inference (e.g., embedding queries forfeature_search)processors/contains your actual logic — model loading, embedding, classification, etc.
Manifest
The manifest is your extractor’s contract with the platform.Vector Index Keys
Multiple vectors are supported — add one entry per embedding your extractor produces.
Inference Type
Declare what kind of real-time inference your extractor provides by settinginference_type in metadata. This lets retriever stages validate that an extractor is compatible with the stage slot.
If
inference_type is omitted, the extractor can only be called via the raw inference endpoint.
Compute Profile
Control resource allocation by addingcompute_profile to your manifest:
BYO Container Image
If your extractor needs native binaries or system packages, specify a custom container image:Batch Processor
Your processor receives a pandas DataFrame and returns it with new columns added.DataFrame Columns
Your__call__ receives these columns:
Batched Processing
Process all rows together — never call a model or API inside a per-row loop:Loading Assets from S3
For binary blobs (images, video, audio), thedata column contains S3 URLs. Use the Extractor SDK to download them:
Pipeline
Wire your processor into the Ray Data pipeline:Row Conditions
Filter which rows a step processes:Using Built-in Models
Compose existing Mixpeek services instead of loading models yourself:
For HuggingFace and your own fine-tuned weights, see the Model Registry.
Real-time Endpoint
Addrealtime.py to expose an HTTP endpoint for query-time inference. This is what lets retriever feature_search stages embed queries with your model.
realtime.py handles embedding only, not retrieval. If your pipeline needs retrieval context (comparing against stored references), configure retriever stages to handle that logic. The real-time endpoint serves on dedicated infrastructure (see Availability).Platform Services
LLM Access
Usecontainer.llm to call platform-managed LLMs with built-in cost tracking and caching:
Secrets
Access encrypted org secrets at runtime viacontainer.secrets:
container.llm instead — it handles API keys automatically.
CLI Tools
Custom extractors can’t importsubprocess directly. Use run_tool for whitelisted CLI tools:
ffmpeg, ffprobe, convert, identify, magick, exiftool, mediainfo, sox, soxi, REDline, art-cmd
Pre-installed Tools
The engine runtime includes these media tools, available viarun_tool:
RED R3D and ARRI RAW decode examples:
Typed SDK
The Extractor SDK provides typed base classes that replace bare Python variables with validated, IDE-friendly types — autocompletion, validation at upload time, and output column checking in the test harness. Import fromshared.extractors.sdk or shared.extractors.
setup() runs once on first batch (lazy model loading); process() runs on each batch. prefetch_hf_model() pre-downloads the model to the HF cache to reduce cold start.
SDK Reference
Security Rules
Custom extractors are scanned before deployment. Code violating these rules is rejected.- Allowed:
numpy,pandas,torch,transformers,sentence_transformers,onnxruntime,PIL,cv2,requests,httpx,os(safe functions only),json,re,pydantic,logging,getattr,hasattr - Blocked:
subprocess,os.system,os.popen,os.exec*,eval,exec,ctypes,socket,multiprocessing,open,setattr,delattr
import os is allowed — only dangerous functions are blocked. Library-internal file I/O (torch.load, transformers.from_pretrained, pd.read_csv) is fine since the scanner only inspects your extractor’s source code.Local Development
Validate and test your extractor locally with the CLI before uploading. No API key needed — these run fully offline and work on any plan.lint catches common mistakes before upload:
- Wrong feature key names (
nameinstead offeature_name) - Missing required fields
- Security scanner violations
test runs your processor through real Ray Data map_batches with Arrow serialization — the same path used in production. Sample rows are fed in the data column (matching production), and the harness validates that your output columns match the manifest features.
Version Management
On a dedicated deployment, the same CLI manages deployed versions with a git-like workflow:
The CLI reads
MIXPEEK_API_KEY, MIXPEEK_NAMESPACE, and MIXPEEK_API_URL from the environment (or --api-key / --namespace).
Archive Limits
Don’t bundle model weights — download from HuggingFace Hub at init time, or use the Model Registry for custom weights.
End-to-End: Extractor → Collection → Retriever
This walkthrough connects all the pieces. SetMIXPEEK_API_KEY, MIXPEEK_NAMESPACE, and MIXPEEK_API_URL first — the same variables the plugins.py CLI reads.
Steps 1 and 5 (the extractor upload/deploy) require a dedicated deployment. On the shared API, use a built-in extractor (e.g.
text_extractor) for the feature_extractor in step 3 and skip steps 1 and 5.1. Deploy the Extractor (dedicated infra)
The simplest path is the CLI —python server/scripts/api/plugins.py push then deploy. The equivalent raw HTTP (custom extractors are addressed as /plugins on this surface; see the Custom Extractor API reference):
2. Create a Bucket and Upload Data
Buckets, collections, retrievers, and batches are top-level resources addressed by the
X-Namespace header — not nested under /namespaces/{ns}/. (Only extractors and models are path-scoped.)3. Create a Collection with the Extractor
Every object uploaded to the bucket flows through this extractor’s batch processor.4. Process the Data (two-step batch)
A batch is created (with the objects + target collections) and then submitted:5. Build a Retriever and Search
At query time the retriever calls your extractor’srealtime.py (dedicated infra) to embed the query, then searches the vectors your batch processor produced.
Troubleshooting
Upload/deploy endpoints return 404 or 405
Upload/deploy endpoints return 404 or 405
The upload/deploy/realtime lifecycle is only available on a dedicated deployment — the shared
api.mixpeek.com exposes only GET list/details for extractors. Develop + lint + test locally, then either provision a dedicated deployment or ship via Submissions.Task COMPLETED but 0 documents
Task COMPLETED but 0 documents
Usually wrong
features key names in manifest.py. Run python server/scripts/api/plugins.py lint path/to/my_extractor to validate. Check that your collection has non-empty vector_indexes via GET /v1/collections/{id}. See Vector Index Keys.Batch produces 0 documents
Batch produces 0 documents
Most common cause: reading from the wrong column. Always use
batch["data"], not batch["text"] or other property names. Check Ray logs for [FailureAggregator] entries.Extractor validation failed
Extractor validation failed
Check
validation_errors. Common issues: using subprocess (use run_tool), using open() directly (use library I/O), using eval/exec (use json.loads).Model loading is slow
Model loading is slow
Use
prefetch_hf_model() in your setup() method to pre-download models to the HF cache. On GKE, HF_HOME points to a shared PVC so models persist across pod restarts. First cold start downloads (~1-2 min); subsequent starts use cache.First run is slow / first search returns 0 (cold start)
First run is slow / first search returns 0 (cold start)
On a cold engine, the embedding model loads on demand — the first batch can take several minutes to leave
PROCESSING, and the first retriever execute afterward may return 0 results (status completed or degraded) while the query-side model warms. This is a cold-start artifact, not a real no-match: retry after a few seconds and results appear. A warm namespace responds immediately. Keep a namespace warm by issuing a periodic lightweight query, or contact your account team about a warm replica floor for latency-sensitive workloads.feature_uri without realtime.py
feature_uri without realtime.py
If your manifest includes a
feature_uri, the system expects a corresponding realtime.py. Without it, feature_search queries against that URI will fail. Omit both if you only need batch processing.Next Steps
Quickstart
Build, test, and query a minimal text embedding extractor end-to-end.
Model Registry
Load HuggingFace models or your own fine-tuned weights inside an extractor.
Extractor Submissions
Submit your extractor for review to be merged into the built-in catalog.
Multi-Tier Extraction
Chain collections into a DAG — transcribe, then embed, then classify.
Reprocess Existing Content
Run a new extractor over an already-ingested corpus — scoped, cost-safe, priced before you run.

