Specifications
Every component, what it is for, and how it extends
The full Mixpeek Appliance build sheet by phase. Each row says what the part does, where it sits, and what it grows into, so nothing is bought twice.
These are planning ranges, not quotes. Hardware pricing, availability and qualification status all move. Validate the seller, the warranty, the qualification status and the interface before buying anything, and treat totals as a way to size a budget rather than as an offer.
Where the budget goes
Each bar sums the low end of the priced line items in that phase, so it is a floor rather than a forecast. The full table is below, unchanged.
Largest: Storage Pod 01: 4U 24-bay ($4,000) · Seagate Exos X24 24 TB SATA HDD x4 ($1,400)
Largest: RTX PRO 6000 Blackwell 96 GB ($13,000) · Threadripper PRO 9975WX ($4,000) · + 8 × 20-24 TB enterprise SATA HDDs ($2,800)
Largest: Lustre shared hot tier ($5,000)
Largest: Dense Seagate-class JBOD/storage shelf ($50,000)
Largest: 48 V LiFePO4 bank ($20,000) · PNW solar array ($15,000) · Hybrid inverter + generator ($8,000)
Bars are the sum of the LOW end of each priced line item, multiplied by the starting quantity where a line is priced per unit, so they are a floor and not an estimate. 7 of 33 rows contribute nothing to them: four carry no figure at all (workload dependent, service dependent, hardware plus service, future node pricing), two are $0 software, and one is an order-of-magnitude entry rather than a number. All seven still appear in the table below. Nothing here is summed into a single figure for the whole build.
| Phase | Component | Purpose | Extends to | Planning range |
|---|---|---|---|---|
| NOW | Storage Pod 01: 4U 24-bayHot-swap server with ECC, BMC, HBA/backplane, redundant PSU; the literal chassis that later slides into the rack, bought now instead of a throwaway enclosure | Makes durable storage independent from compute from day one; never gets replaced | Populate 24 bays over time; Storage-02/N; dense JBOD later | $4,000-$8,000 before disks |
| NOW | Seagate Exos X24 24 TB SATA HDDEnterprise 3.5-inch SATA disk | Populate Storage-01 from day one; drives never migrate because the chassis never changes | 4 drives = 96 TB raw now; fill toward 24 bays over time, same chassis through V1 | $350-$500 each |
| NOW | ZFS + S3-compatible object serviceDurable implementation + stable object API | Canonical `s3://` namespace live from day one, not created during a later migration | Backend can become distributed object storage | $0+ software |
| NOW | Independent backupSeparate NAS/disks/cloud copy | Avoid a single-copy event | Grows with corpus | Workload dependent |
| V1 | 27-32U enclosed vertical rackFull-depth, lockable, wheeled, high-airflow cabinet | Clean movable mini data center | Add pods until full; duplicate rack later | $1,000-$2,500 |
| V1 | Compute Pod 01: 4U chassisFull-depth GPU chassis; server airflow/serviceability preferred | Replaceable AI compute unit | GPU #2 vertically; Compute-02/N horizontally | $700-$2,000 |
| V1 | RTX PRO 6000 Blackwell 96 GB96 GB ECC NVIDIA accelerator; Server Edition preferred in qualified server | Large coding/VLM/video inference, Mixpeek extraction, training/fine-tuning | Second 96 GB GPU; future RTX PRO/MGX/HGX/DGX nodes | $13,000-$15,000 |
| V1 | Threadripper PRO 9975WX32C/64T, high PCIe capacity | FFmpeg, Ray CPU workers, preprocessing, tokenization, DBs, GPU feeding | Add CPU/GPU nodes; higher-core node if measured need | $4,000-$5,000 |
| V1 | ASUS Pro WS WRX90E-SAGE SEECC RDIMM, PCIe 5.0, dual 10GbE, AST2600 BMC/IPMI | Expansion + RAM bandwidth + remote recovery | Reserve x16 for GPU #2 and ConnectX NIC | $1,200-$1,500 |
| V1 | 512 GB ECC DDR5 RDIMMRegistered ECC memory | Dataset staging, Ray, dataloaders, CPU offload, training | Scale toward ~2 TB; each node adds RAM | $1,500-$3,000 |
| V1 | 2 × 2 TB mirrored boot NVMeRedundant system volume | Ubuntu, configs, critical service state | Larger mirror/dedicated management storage later | $300-$600 |
| V1 | 8 TB model/dataset NVMeFast reusable local tier | Weights, HF cache, tokenized/hot datasets | Add/larger enterprise NVMe per node | $700-$1,500 |
| V1 | 4-8 TB high-endurance scratch NVMeSeparate high-write tier | Frames/audio, checkpoints, optimizer state, temp tensors | Add NVMe; shared hot data moves to Lustre at V2 | $500-$1,500 |
| V1 | 1600 W+ PSU / OEM redundant PSUGPU-capable power | Stable compute power with headroom | GPU #2 where envelope permits | $500-$1,000 |
| V1 | Cooling/fansTR5/server cooling, front-to-back airflow | Sustained 24/7 operation | OEM/liquid cooling for dense future accelerators | $250-$700 |
| V1 | + 8 × 20-24 TB enterprise SATA HDDsGrows the pool already running since NOW, to ~12 drives total | Canonical media, Mixpeek objects, datasets, artifacts | ~240-288 TB raw at ~12 drives; ~480-576 TB with all 24 bays full; then PB shelves | $2,800-$4,000 |
| V1 | 100 GbE ConnectX-class NICsHigh-speed node adapters | Keep storage traffic from starving GPUs | 100 → 200/400 GbE RDMA | $500-$2,000/node |
| V1 | 100 GbE managed data switchInternal high-speed fabric | Compute ↔ storage and future shared tier | Upgrade to 200/400G; split fabrics later | $1,500-$5,000 |
| V1 | Management switch/VLANSeparate OOB network | BMC/IPMI, UPS, switches, controllers, sensors | Add all future nodes/racks | $200-$750 |
| V1 | Kubernetes control nodeSmall dedicated x86 host | Keeps orchestration independent of GPU workers | 1 → 3 HA controllers | $300-$800 |
| V1 | Local consoleSmall monitor + keyboard/mouse + shelf | Break-glass install/BIOS/network recovery | Rack KVM later | $150-$500 |
| V1 | Rack PDU + UPSMetered power + short ride-through | Clean power and graceful shutdown | Larger/redundant UPS as load grows | $1,500-$4,000 |
| V1 | Rails/optics/DAC/SAS/power/sparesIntegration hardware | Repeatable, serviceable installation | Standardize for future pods | $500-$1,500 |
| V1 | Tailscale access planeSecure overlay for clients/admin | Mac/amux and engineers reach services without rack depending on Mac | Add users/sites; substitute approved on-prem solution for true air-gap | Service dependent |
| V2 | Lustre shared hot tierParallel POSIX filesystem over fast/RDMA fabric | Shared training/video working set; avoids node-local data dependency | Add MDS/OSS/OST capacity | $5,000-$25,000+ |
| V2 | Compute Pod 02Second GPU worker node | Horizontal throughput + distributed training | Compute-03/N | Future node pricing |
| V2 | Kubernetes + GPU Operator + KubeRayCluster scheduling/distributed execution | One logical resource pool | Add workers horizontally | $0+ software |
| FUTURE | 200/400 GbE or InfiniBand compute fabricTightly coupled inter-node network | Distributed training / giant model parallelism | Future HGX/DGX/MGX nodes | $10Ks+ |
| FUTURE | Dense Seagate-class JBOD/storage shelf60-106-drive expansion | Hundreds of TB → PB-scale object storage | Add shelves independently of compute | $50K-$150K+ |
| OFF-GRID | Starlink + secure accessRemote WAN | Cabin connectivity | Backup WAN | Hardware + service |
| OFF-GRID | 48 V LiFePO4 bankLarge-format battery storage | Overnight/low-solar operation | Add modules from measured kWh/day | ~$20K-$35K+ V1 scale |
| OFF-GRID | PNW solar arrayGround array; planning target ~25-30 kW for year-round V1 resilience | Generate rack energy + recharge batteries | Expand from measured load/site yield | ~$15K-$40K+ installed |
| OFF-GRID | Hybrid inverter + generatorAC conversion/charging + long-dark-period backup | Resilient off-grid operation | Parallel/larger equipment as compute grows | $8K-$25K+ |
Showing 33 of 33 components.
Worked examples
What the sheet above actually buys, in capacity terms. Every figure is arithmetic from the assumption printed beside it, so you can substitute your own bitrate or drive size and redo it. None of these are benchmarks: we publish prices and capacities, and we are not going to publish throughput numbers we have not measured.
How much 1080p video fits in the V1 storage pod?
Assumptions
- ·12 x 24 TB enterprise SATA, the V1 starting population
- ·Two 6-wide RAIDZ2 vdevs, so 4 of the 12 drives are parity
- ·1080p H.264 at roughly 2 GB per hour
Roughly 86,000 hours of 1080p video, about 9.9 years of continuous footage.
Swap the bitrate for your own and the division is the only thing that changes. 4K at 8 GB per hour lands nearer 21,000 hours.
What happens when you fill all 24 bays?
Assumptions
- ·The same pod, 24 bays populated at 24 TB
- ·Four 6-wide RAIDZ2 vdevs, the same layout repeated
The pod roughly doubles without a second chassis, a second rack, or any change to the object namespace.
This is the point of buying the 24-bay chassis at V1 rather than a 12-bay one. The expansion is drives, not architecture.
What does Phase 0 buy before the rack exists?
Assumptions
- ·The final 24-bay storage server, bought now rather than a desktop enclosure
- ·Populated with 4 x 24 TB enterprise SATA to start
- ·Runs beside the desk on ordinary 1/10 GbE until the rack cabinet exists
96 TB of usable corpus beside the desk now, on the exact chassis that later slides into the rack. There is no migration event, because nothing moves.
Mirrored or parity layouts cut the usable figure; this is the raw number. It costs more than the desktop-enclosure plan it replaced, and what the extra buys is that nothing here is ever replaced and the canonical s3:// namespace exists from day one instead of being created during a later migration.
Software stack
All of it runs in the rack. Nothing in this list depends on a hosted service to function.
- Mixpeek local stack
- Ray / KubeRay
- Kubernetes
- NVIDIA GPU Operator
- vLLM
- SGLang
- TensorRT-LLM
- NeMo
- Megatron
- PyTorch
- Hugging Face
- Qdrant
- S3-compatible object service
- ZFS
- Lustre (V2)
- Prometheus / Grafana-class observability
The one rule behind the whole sheet
Every NOW-phase line is chosen so it is not thrown away at V1, and the plan is stricter than that now: the storage server bought in Phase 0 is the same physical unit that later slides into the rack, so there is no migration event to survive. It costs more up front than a desktop enclosure would, and what that buys is that the canonical s3:// namespace exists from the first day rather than being created during a later copy-and-verify. That constraint is why the build order starts with storage: it is the only layer where buying early does not mean buying twice.
Not sure what you would actually need?
Tell us the corpus, the workloads and the constraints, and we will size it against the build sheet: drive count, GPU class, fabric, power, and which phase to start at. You own the hardware either way, and the plan is published in full whether you build it yourself or have us do it.
It is fully S3-compatible, with no code change. The object endpoint speaks the S3 API, so anything already talking to S3 points at a new endpoint URL and keeps working. Same SDKs, same bucket and key layout, same tooling.