Skip to main content
POST
Trigger Collection Processing

Authorizations

Authorization
string
header
required

Mixpeek API key, sent as Authorization: Bearer mxp_sk_.... Create one in Studio under Settings → API Keys, or with an admin key via POST /v1/organizations/users/{user_email}/api-keys. A missing header returns 403; an invalid or revoked key returns 401.

X-Namespace
string
header
required

Namespace id (ns_...), not the namespace name. This scopes the request rather than authenticating it, and it is required on every operation marked x-mixpeek-namespace-scoped.

Path Parameters

collection_identifier
string
required

The ID or name of the collection to trigger

Body

application/json

Request to trigger (re)processing through a collection.

For bucket-sourced collections (tier 0): Discovers objects from source bucket(s) and creates a batch for processing. Use include_buckets to limit which source buckets to process from.

For collection-sourced collections (tier N): Processes existing documents from upstream collection(s). Use include_collections to limit which source collections to process from.

Use source_filters for field-level filtering on objects or documents.

Document Overwrite Behavior:

  • If source bucket has unique_key configured: Documents are UPSERTED (overwrites existing)
  • If source bucket has NO unique_key: New documents are CREATED (may cause duplicates)

To enable idempotent re-processing, configure unique_key on the source bucket.

include_buckets
string[] | null

Limit processing to objects from these specific buckets (IDs or names). Only applies to bucket-sourced collections. If not provided, all configured source buckets are used.

include_collections
string[] | null

Limit processing to documents from these specific collections (IDs or names). Only applies to collection-sourced collections. If not provided, all configured source collections are used.

object_ids
string[] | null

Limit processing to these specific object IDs. Only applies to bucket-sourced collections. This is a convenience shorthand — equivalent to using source_filters with {"AND": [{"field": "object_id", "operator": "in", "value": [...]}]}.

source_filters
LogicalOperator · object | null

Field-level filters for objects (bucket-sourced) or documents (collection-sourced). Uses LogicalOperator format (AND/OR/NOT). Use this to filter by metadata fields, status, or any other object/document properties.

Example:
dedup_strategy
enum<string> | null

How to handle sources already processed in prior batches. skip (default): skip sources already materialized in this collection. replace: delete existing documents for the re-processed sources and re-materialize them — this also clears the processed-objects resume ledger, so use it to recover a collection stuck with ledger entries but 0 materialized documents (the orphan/divergence state). force: process regardless, allowing duplicates.

Available options:
skip,
replace,
force
reembed
StoredTextReembedRequest · object | null

Re-embed stored document text instead of reprocessing sources. When set, no batch is created: the collection's existing documents are read, the selected embedding outputs are recomputed from the text each document already stores, and the vectors are written back onto the same documents. Progress is reported on the returned task. Source selection fields (include_buckets, include_collections, object_ids, source_filters, dedup_strategy) cannot be combined with it.

Example:

Response

Successful Response

Response after triggering collection processing.

Use batch_id or task_id to monitor progress via GET /v1/batches/{batch_id} or GET /v1/tasks/{task_id}.

task_id
string
required

Task ID for monitoring via GET /v1/tasks/{task_id}.

collection_id
string
required

ID of the collection being processed.

total_tiers
integer
required

Number of processing tiers in the DAG.

message
string
required

Human-readable status message.

batch_id
string | null

ID of the created batch for tracking progress. Null for a stored-text re-embed (reembed), which runs as a task, not a batch.

source_bucket_ids
string[] | null

Bucket IDs that objects were discovered from (bucket-sourced collections).

source_collection_ids
string[] | null

Collection IDs that documents were read from (collection-sourced collections).

object_count
integer | null

Total number of objects included in the batch (bucket-sourced collections).

batch_ids
string[] | null

All batch IDs created for this trigger. Present when a large load was auto-chunked into multiple right-sized batches; batch_id is the first of these. Null (absent) for a single-batch trigger.

task_ids
string[] | null

Task IDs for each created batch, aligned with batch_ids.

batch_count
integer | null

Number of batches created. Null for a single-batch trigger; set when the load was auto-chunked into multiple batches.

document_count
integer | null

Total number of documents to process (collection-sourced collections).