Skip to main content
Mixpeek is split into two deployable components:
  • API Layer – FastAPI + Celery + Redis connection (HTTP endpoints, task orchestration, webhooks).
  • Engine Layer – Ray cluster + Ray Serve (extractors, inference, clustering, taxonomy runs).
Shared dependencies: MongoDB, MVS, Redis, and S3-compatible object storage.

Local Development

./start.sh scripts spin up a full stack with Docker Compose:
Docker Compose services:
  • mongodb – metadata (mongodb://localhost:27017)
  • mvs – vector storage (MVS)
  • redis – task queue/cache (redis://localhost:6379)
  • localstack – S3 emulator (http://localhost:4566)
Run curl http://localhost:8000/v1/health to confirm readiness, then follow the Quickstart.

Production Topology (Kubernetes)

Recommended node pools:
  • API nodes – general purpose (e.g., t3.xlarge), scale FastAPI/Celery horizontally.
  • CPU workers – compute-optimized (e.g., c5.4xlarge) for text extraction, clustering.
  • GPU workers – GPU instances (e.g., p3.2xlarge) for embeddings, rerankers, video processing.
Expose the API via an ingress or load balancer; keep Ray Serve internal unless exposing custom inference endpoints.

Ray on Kubernetes (KubeRay)

This page describes Mixpeek operating a Ray cluster it deploys. If you already run Ray or KubeRay and want Mixpeek to use yours instead, see Bring Your Own Ray — it covers the three contracts, every permission we ask for, and what we check before deploying anything.
The Engine layer runs as a RayService, a custom resource the KubeRay operator reconciles. You apply one YAML; KubeRay creates the head pod, the worker groups, and the services in front of them, and replaces them on config change. You do not create Ray pods yourself, and there is nothing to scale by hand. Each worker group carries its own minReplicas and maxReplicas, and Ray’s autoscaler moves within that range on pending-task pressure. The API layer reaches the cluster over ENGINE_API_URL, a plain in-cluster HTTP address such as http://mixpeek-engine-svc-head-svc.mixpeek-engine:8265.
Job submission carries no credential. The API builds a bare http://host:port to the Ray dashboard and submits against it, so reachability is the whole access control. Keep the Ray dashboard port off any public ingress and restrict it with a NetworkPolicy. Anything that can reach the port can submit a job.

Namespace-scoped or cluster-wide

Two different scopes, and only one of them is per-deployment: So a second Mixpeek deployment on the same cluster means a second namespace with its own RayService, not a second operator. The operator needs cluster-level permission to create pods and services in the namespaces it watches, which is the one privilege you cannot scope away. Isolate deployments by namespace, one per tenant, with a LimitRange and a NetworkPolicy on each. Separate node pools with taints keep GPU work off shared capacity, since namespaces alone do not partition nodes.

Core Environment Variables

Secrets should be injected via Kubernetes secrets, environment managers, or cloud secret stores.

Health & Verification

  • Endpoint: GET /v1/health – checks Redis, MongoDB, MVS, Celery, Engine, ClickHouse (if enabled).
  • Smoke test: create namespace → bucket → collection → upload object → submit batch → execute retriever.
  • Tasks: ensure Celery workers process webhook events, cache invalidations, and maintenance tasks.

Scaling Guidelines

Monitor Ray dashboard (port 8265) for job status, resource utilization, and Serve deployments.

Deployment Checklist

  1. Provision MongoDB, MVS, Redis, and S3/GCS buckets (with IAM roles).
  2. Deploy Ray cluster (head + workers) and confirm job submission works.
  3. Deploy FastAPI + Celery services; configure environment variables to point to Ray + data stores.
  4. Configure ingress/HTTPS, secrets, and network policies.
  5. Run health checks and quickstart workflow to verify end-to-end functionality.
  6. Set up observability (logs, metrics, webhooks) and configure backups for MongoDB.

References