Skip to main content

Browse the extractor catalog on GitHub

Runnable reference for every built-in Mixpeek extractor — inputs, parameters, output fields, embedding models, and copy-paste examples. Auto-generated from the live registry, so it always matches production.
Custom extractor flow: a single extractor archive powers both the ingestion pipeline (batch Ray Data) and the retrieval realtime endpoint (Ray Serve HTTP), used by retriever stages like feature_search, rerank, and agentic_enrich
Custom extractors let you run your own code on Mixpeek infrastructure — inside the same Ray cluster that powers the built-in extractors and retriever stages. You keep full control of the logic, model, and I/O; Mixpeek handles packaging, scheduling, GPU allocation, caching, and observability.

Availability

The upload → deploy → real-time lifecycle runs on dedicated Enterprise infrastructure — it is not exposed on the shared public API at api.mixpeek.com. What works where:On the shared API the upload/deploy/realtime endpoints return 404/405 by design. Contact your account team to provision a dedicated deployment for self-service custom-extractor uploads, or use Submissions to ship an extractor into the built-in catalog. The dedicated upload/deploy/realtime HTTP contract is documented in the Custom Extractor API (Dedicated Infrastructure) reference.
Want your extractor available to everyone without dedicated infra? Submit it for review to be merged into the built-in catalog — see Extractor Submissions.

What You Can Build

Custom extractors plug into two places in the warehouse:

1. Feature Extractors (Decomposition)

Attach custom logic to a collection so every ingested object flows through your pipeline during decomposition. Your extractor is your vocabulary — unlike built-in pipelines (which are selected via features), custom extractors keep their explicit names and are selected as a custom feature:
The explicit feature_extractor config (shown later on this page) also works — use it when you need custom input_mappings or parameters. Use custom extractors to:
  • Embed domain-specific content with your own model (fine-tuned CLIP, proprietary audio encoder, etc.)
  • Extract structured attributes via a VLM you manage (brand compliance, regulated content classification)
  • Transcribe, OCR, or segment media with a custom pre/post-processing chain
  • Produce multiple named vector indexes from a single pass
Outputs land in MVS and MongoDB with the same feature URI scheme as built-in pipelines (mixpeek://my_extractor@1.0.0/my_embedding), so retrievers, taxonomies, and clusters can reference them.
Pricing: a custom:<name> feature derives its per-unit rate from the compute profile your extractor declares — the same machinery that prices native features. Quote it before running via POST /v1/organizations/billing/estimate; see Billing & Pricing.

2. Retriever Operations (Query Time)

An extractor’s realtime.py exposes a Ray Serve HTTP endpoint that retriever stages can call during execution. Use this to power:
  • feature_search — embed queries at search time with the same model you used during ingestion, so the query vector lives in the same space as the indexed vectors
  • Inference operations — on-the-fly classification, scoring, or re-ranking against your model
  • LLM calls — wrap a hosted or private LLM behind a stable contract, with platform-managed secrets and cost tracking
  • Classifiers — apply your own classifier to candidate results mid-pipeline
This is what lets a single custom extractor own both halves of a retrieval flow: it encodes documents on the way in, and encodes the query on the way out.

Delivery Formats

Ship your extractor in either of two formats: Both formats expose the same runtime APIs (batch __call__, real-time run_inference, platform LLM/secret accessors).

Extractor Structure

Every extractor has the same layout:
  • manifest.py declares what your extractor accepts, produces, and which vector indexes to create
  • pipeline.py wires your processor into the Ray Data batch pipeline
  • realtime.py exposes a Ray Serve endpoint for query-time inference (e.g., embedding queries for feature_search)
  • processors/ contains your actual logic — model loading, embedding, classification, etc.

Manifest

The manifest is your extractor’s contract with the platform.

Vector Index Keys

Use the exact key names below. Wrong keys silently create a collection with no vector indexes — your batch will show COMPLETED but produce 0 documents.
Multiple vectors are supported — add one entry per embedding your extractor produces.

Inference Type

Declare what kind of real-time inference your extractor provides by setting inference_type in metadata. This lets retriever stages validate that an extractor is compatible with the stage slot. If inference_type is omitted, the extractor can only be called via the raw inference endpoint.

Compute Profile

Control resource allocation by adding compute_profile to your manifest:
For API-based or hash-based extractors that don’t need GPU, set resource_type: "cpu" to skip GPU allocation — saves ~3 minutes startup time and costs ~6x less.

BYO Container Image

If your extractor needs native binaries or system packages, specify a custom container image:
Base your image on the Mixpeek engine image to get Ray, FFmpeg, and all SDK helpers:
Images must be pushed to your org-scoped Artifact Registry repo. GKE Workload Identity handles pull auth. Contact your account team to provision access.

Batch Processor

Your processor receives a pandas DataFrame and returns it with new columns added.

DataFrame Columns

Your __call__ receives these columns:
Always read from batch["data"] — not a column named after your blob property. If your bucket has a text property, the content is still in batch["data"]. (The local test harness feeds the same data column, so an extractor that passes test behaves the same in production.)

Batched Processing

Process all rows together — never call a model or API inside a per-row loop:

Loading Assets from S3

For binary blobs (images, video, audio), the data column contains S3 URLs. Use the Extractor SDK to download them:

Pipeline

Wire your processor into the Ray Data pipeline:

Row Conditions

Filter which rows a step processes:

Using Built-in Models

Compose existing Mixpeek services instead of loading models yourself:
For HuggingFace and your own fine-tuned weights, see the Model Registry.

Real-time Endpoint

Add realtime.py to expose an HTTP endpoint for query-time inference. This is what lets retriever feature_search stages embed queries with your model.
The return dict must include an "embedding" key — this is what feature_search uses as the query vector. For multi-vector extractors, include additional keys matching your feature_name values.
realtime.py handles embedding only, not retrieval. If your pipeline needs retrieval context (comparing against stored references), configure retriever stages to handle that logic. The real-time endpoint serves on dedicated infrastructure (see Availability).

Platform Services

LLM Access

Use container.llm to call platform-managed LLMs with built-in cost tracking and caching:

Secrets

Access encrypted org secrets at runtime via container.secrets:
For platform LLMs, use container.llm instead — it handles API keys automatically.

CLI Tools

Custom extractors can’t import subprocess directly. Use run_tool for whitelisted CLI tools:
Available tools: ffmpeg, ffprobe, convert, identify, magick, exiftool, mediainfo, sox, soxi, REDline, art-cmd

Pre-installed Tools

The engine runtime includes these media tools, available via run_tool: RED R3D and ARRI RAW decode examples:

Typed SDK

The Extractor SDK provides typed base classes that replace bare Python variables with validated, IDE-friendly types — autocompletion, validation at upload time, and output column checking in the test harness. Import from shared.extractors.sdk or shared.extractors.
setup() runs once on first batch (lazy model loading); process() runs on each batch. prefetch_hf_model() pre-downloads the model to the HF cache to reduce cold start.

SDK Reference


Security Rules

Custom extractors are scanned before deployment. Code violating these rules is rejected.
  • Allowed: numpy, pandas, torch, transformers, sentence_transformers, onnxruntime, PIL, cv2, requests, httpx, os (safe functions only), json, re, pydantic, logging, getattr, hasattr
  • Blocked: subprocess, os.system, os.popen, os.exec*, eval, exec, ctypes, socket, multiprocessing, open, setattr, delattr
import os is allowed — only dangerous functions are blocked. Library-internal file I/O (torch.load, transformers.from_pretrained, pd.read_csv) is fine since the scanner only inspects your extractor’s source code.

Local Development

Validate and test your extractor locally with the CLI before uploading. No API key needed — these run fully offline and work on any plan.
lint catches common mistakes before upload:
  • Wrong feature key names (name instead of feature_name)
  • Missing required fields
  • Security scanner violations
test runs your processor through real Ray Data map_batches with Arrow serialization — the same path used in production. Sample rows are fed in the data column (matching production), and the harness validates that your output columns match the manifest features.

Version Management

On a dedicated deployment, the same CLI manages deployed versions with a git-like workflow: The CLI reads MIXPEEK_API_KEY, MIXPEEK_NAMESPACE, and MIXPEEK_API_URL from the environment (or --api-key / --namespace).

Archive Limits

Don’t bundle model weights — download from HuggingFace Hub at init time, or use the Model Registry for custom weights.

End-to-End: Extractor → Collection → Retriever

This walkthrough connects all the pieces. Set MIXPEEK_API_KEY, MIXPEEK_NAMESPACE, and MIXPEEK_API_URL first — the same variables the plugins.py CLI reads.
Steps 1 and 5 (the extractor upload/deploy) require a dedicated deployment. On the shared API, use a built-in extractor (e.g. text_extractor) for the feature_extractor in step 3 and skip steps 1 and 5.

1. Deploy the Extractor (dedicated infra)

The simplest path is the CLI — python server/scripts/api/plugins.py push then deploy. The equivalent raw HTTP (custom extractors are addressed as /plugins on this surface; see the Custom Extractor API reference):

2. Create a Bucket and Upload Data

Buckets, collections, retrievers, and batches are top-level resources addressed by the X-Namespace header — not nested under /namespaces/{ns}/. (Only extractors and models are path-scoped.)

3. Create a Collection with the Extractor

Every object uploaded to the bucket flows through this extractor’s batch processor.

4. Process the Data (two-step batch)

A batch is created (with the objects + target collections) and then submitted:
At query time the retriever calls your extractor’s realtime.py (dedicated infra) to embed the query, then searches the vectors your batch processor produced.

Troubleshooting

The upload/deploy/realtime lifecycle is only available on a dedicated deployment — the shared api.mixpeek.com exposes only GET list/details for extractors. Develop + lint + test locally, then either provision a dedicated deployment or ship via Submissions.
Usually wrong features key names in manifest.py. Run python server/scripts/api/plugins.py lint path/to/my_extractor to validate. Check that your collection has non-empty vector_indexes via GET /v1/collections/{id}. See Vector Index Keys.
Most common cause: reading from the wrong column. Always use batch["data"], not batch["text"] or other property names. Check Ray logs for [FailureAggregator] entries.
Check validation_errors. Common issues: using subprocess (use run_tool), using open() directly (use library I/O), using eval/exec (use json.loads).
Use prefetch_hf_model() in your setup() method to pre-download models to the HF cache. On GKE, HF_HOME points to a shared PVC so models persist across pod restarts. First cold start downloads (~1-2 min); subsequent starts use cache.
On a cold engine, the embedding model loads on demand — the first batch can take several minutes to leave PROCESSING, and the first retriever execute afterward may return 0 results (status completed or degraded) while the query-side model warms. This is a cold-start artifact, not a real no-match: retry after a few seconds and results appear. A warm namespace responds immediately. Keep a namespace warm by issuing a periodic lightweight query, or contact your account team about a warm replica floor for latency-sensitive workloads.
If your manifest includes a feature_uri, the system expects a corresponding realtime.py. Without it, feature_search queries against that URI will fail. Omit both if you only need batch processing.

Next Steps

Quickstart

Build, test, and query a minimal text embedding extractor end-to-end.

Model Registry

Load HuggingFace models or your own fine-tuned weights inside an extractor.

Extractor Submissions

Submit your extractor for review to be merged into the built-in catalog.

Multi-Tier Extraction

Chain collections into a DAG — transcribe, then embed, then classify.

Reprocess Existing Content

Run a new extractor over an already-ingested corpus — scoped, cost-safe, priced before you run.