Batch Processing
Multi-tier DAG processing that transforms bucket objects into searchable documents with feature extraction
Why do anything?
Raw objects need feature extraction to become searchable. Without batch processing, you can't generate embeddings or enrich content at scale.
Why now?
AI search requires vector embeddings. Manual processing doesn't scale.
Why this feature?
Multi-tier DAG processing handles complex pipelines: Tier 0 (bucket→collection) and Tier N (collection→collection) with Celery workers and Ray inference.
How It Works
Batch processing uses a multi-tier DAG architecture. Tier 0 processes bucket objects, Tier N processes upstream collection outputs.
Batch Creation
Create batch record, validate source and collection config
Task Routing
Route to Celery process_tier queue based on tier level
Feature Extraction
Ray engine runs feature extractor, generates embeddings
Document Storage
Documents stored in the Mixpeek Vector Store with vectors indexed
Why This Approach
DAG enables complex multi-stage pipelines (e.g., video→frames→faces→embeddings). Celery provides reliable task execution. Ray handles ML inference at scale.
Where This Is Used
Recent updates
Full changelog- Sep 26, 2026Batch status is right when a tier runs several extractorsWhen one tier ran two extractors, such as passthrough and text on the same bucket, the first to finish could close the tier, and the batch ended as completed with errors saying no documents were written, then the documents arrived a minute later. Each extractor job now reports its own job id, so a tier waits for all of them. A batch that was re-dispatched from the queue also marks its objects completed; they used to stay pending.
- Sep 27, 2026Small CPU batches start soonerA CPU batch sized its workers from the processing profile's maximum rather than from the number of objects, so a one-document batch started several workers and could wait for capacity it never needed. Workers are now sized from the objects in the batch and from what the shared pool can hold.
- Sep 27, 2026Batch diagnostics label out-of-memory failures as infrastructureA batch stopped by Ray's memory monitor was recorded with an unknown failure category. It is now labelled infrastructure, like other out-of-memory failures, so it reads as retryable. Diagnostics for a batch that is still queued now list real next steps.
- Sep 26, 2026Batches of more than 500 objects submit againSubmitting a batch read all of its objects in one query, which went over the 1,000-row page limit, so a batch of 2,000 objects was refused with a 400. Objects are now read 500 at a time.