NEWVectors or files. Pick a path.Start →

    FAQ

    Questions people actually ask before buying a rack

    Including the ones where the honest answer is that you should not buy one.

    Is this the same as AWS S3?
    It speaks the same API, and it is not the same thing. S3 is a service in someone else's region, billed per gigabyte stored and per gigabyte read. This is an S3-compatible object service running on ZFS on disks in your rack, so the API your code targets is unchanged while the bytes, the failure domain and the read cost are all yours. The compatibility is the point: it means the corpus stays readable by anything that speaks S3, including tooling that has nothing to do with Mixpeek.
    Can I still use cloud AI models with this?
    Yes, as an explicit choice per workload rather than a default. Models you self-host run locally on your own accelerator and involve no egress at all. Where a workload genuinely needs a frontier model that cannot be self-hosted, Mixpeek can mediate a controlled path: upload, call, log, delete, recorded so the boundary is auditable. The honest version of this claim is that such a call does send data out, and the difference is that somebody decided to and it was logged.
    What exactly is the Mixpeek Appliance?
    The Mixpeek Appliance is a private mini data center for multimodal AI: a 27-32U rack that holds a GPU compute node, a 24-bay storage node and the networking between them, running the full Mixpeek stack locally. Durable data lives behind an S3-compatible object API on your own disks; Mixpeek orchestrates extraction, indexing and retrieval against that object namespace. It is a private, object-storage-first multimodal AI appliance rather than a server with a GPU in it.
    Why is it object storage first rather than filesystem first?
    Because the durable asset is the corpus and its namespace, and compute is replaceable. Addressing data as s3:// objects rather than as paths on a particular machine means the disks, the filesystem and the servers underneath can all change without any pipeline, dataset reference or index having to change with them. Local NVMe exists as cache and scratch, and nothing treats it as the source of truth.
    Do I have to buy the whole rack to start?
    No, and the build order is designed around that. Phase 0 is enterprise SATA disks in a Thunderbolt enclosure, which is useful storage immediately and is chosen specifically so those disks migrate into the 24-bay storage pod later instead of being thrown away. The rack is the second purchase, not the first.
    How does this extend to a fully on-premise deployment?
    It already is one. Every layer, object storage, extraction, embedding, vector search, inference and training, runs on hardware in the rack. The only outside dependency in the reference build is a secure overlay network for reaching services, and that is a client convenience rather than an infrastructure dependency: swap it for an approved on-premise access solution and the system is air-gapped without any other change.
    What happens when I outgrow one rack?
    Each dimension scales on its own. More inference or training throughput means another compute pod. A larger model means another accelerator. More durable capacity means more disks, another storage pod, or a dense shelf. Shared multi-node throughput means adding the parallel filesystem tier. More bandwidth means moving the fabric from 100 GbE toward 200/400 GbE or InfiniBand. When the rack is full, you add a rack, and the object namespace does not change.
    Does buying this now mean buying the wrong thing before newer hardware ships?
    The design assumes newer hardware. Future qualified accelerator nodes join the same Kubernetes and Ray cluster as additional workers, and because the durable data sits behind a stable object API rather than inside any one node, adding them invalidates nothing. That is the reason the storage layer, not the GPU, is the thing the architecture is organised around.
    How does this work alongside cloud?
    As a hybrid, with the split drawn on read volume. The workloads that read the corpus repeatedly, extraction, re-indexing, training and evaluation, run locally against local objects. Cloud stays available for burst capacity, distribution and anything genuinely global. The reason to draw the line there is that multimodal indexing re-reads everything each time a model or pipeline changes, and that read pattern is what makes remote object storage expensive.
    What software runs on it?
    The local Mixpeek stack for extraction, indexing and multi-stage retrieval; Ray and KubeRay for distributed execution; Kubernetes with the NVIDIA GPU Operator; vLLM, SGLang and TensorRT-LLM for serving; NeMo, Megatron, PyTorch and Hugging Face for training and fine-tuning; a vector service; ZFS behind the S3-compatible object service; and Prometheus/Grafana-class observability.
    Can it run off-grid?
    That is an optional power module rather than part of the core build, and it is sized per site rather than from a catalogue. It combines a solar array, a 48 V LiFePO4 bank, a hybrid inverter and a generator for long dark periods, with environmental telemetry and energy-aware scheduling so the cluster's duty cycle can follow available power. Sizing depends on measured load and site yield, so the figures in the spec sheet are planning ranges rather than a design.
    Are the prices quotes?
    No. Every figure on the specifications page is a planning range taken from the internal build document, and hardware pricing, availability and qualification status all move. Validate the seller, the warranty, the qualification status and the interface before purchasing anything, and treat the totals as a way to size a budget rather than as an offer.

    Mixpeek Appliance vs public cloud, plainly

    Both directions. If your answer is in the right-hand column more often than the left, buy the cloud version and do not build a rack.

    Mixpeek Appliance compared with public cloud object storage plus hosted AI
    DimensionMixpeek AppliancePublic cloud + hosted AI
    Where the data sitsDisks in a rack you control, behind an S3-compatible APIA provider's region, behind the same API shape
    Cost of reading the corpusPower and hardware you already boughtMetered egress and request charges on every re-read
    Cost to startCapital up front, from roughly $2.3K for storage-first to tens of thousands for a V1 rackNear zero, then a bill that grows with what you keep and read
    Time to first resultProcurement, assembly and setupMinutes
    ElasticityFixed until you buy more; you own the peakEffectively unlimited, billed accordingly
    Who fixes a failed diskYou, or whoever you contracted toNobody tells you it happened
    Regulatory postureResidency is a physical factResidency is a contractual assurance plus region selection
    Frontier model accessSelf-hosted by default; external calls are an explicit mediated exceptionNative, and the data goes with the call
    Scaling limitRack space, power and coolingBudget

    Sovereign by default, cloud-capable when you choose it

    Models you self-host, served locally with vLLM, SGLang or TensorRT-LLM on your own accelerator, never leave the building. That is the default and it needs no exception process. Some frontier models cannot be self-hosted at all, and when a workload genuinely needs one, Mixpeek can mediate a controlled egress path rather than leaving you to build one: upload, call, log, delete, as an explicit per-workload opt-in.

    • Self-hosted inference is the default path and involves no egress at all.
    • Calling an external frontier model is an explicit decision per workload, not a global setting.
    • Each mediated call is logged, so the boundary is auditable rather than assumed.
    • Content sent for a mediated call is deleted after the call rather than retained.

    The honest framing is that this is flexibility with a control on it, not a claim that nothing ever leaves. A workload that calls an external model does send data out; the difference is that it is a decision somebody made, recorded, rather than a default nobody noticed.

    Sponsored deployment: you fund the build, Mixpeek deploys and can operate it

    One route to owning a Mixpeek Appliance is to sponsor the hardware build-out and have Mixpeek specify, assemble, deploy and optionally operate it as your private local S3 plus AI instance. The alternative is to build it yourself from the public plan, which is why that plan is published in full rather than gated.

    Who owns the hardware?
    You do. It is your capital purchase, sitting in your facility, and it stays yours regardless of what happens to the software relationship.
    Who is on call when a disk dies?
    That is the part worth negotiating explicitly rather than assuming. Hardware failure response, spares holding and replacement windows are a scoped operational commitment, and the scope should be written down before anyone signs.
    What happens at the end of a contract?
    The hardware and the data are yours and stay where they are. The object namespace is a standard S3-compatible API, so the corpus remains readable by anything that speaks S3 whether or not Mixpeek is still involved.
    Is this managed-service-forever?
    No, and claiming otherwise would be dishonest about what operating physical infrastructure involves. It is a real operational commitment with a defined scope, and the scope is the thing to agree on.

    Prefer to build it yourself? The full plan is published, including the bill of materials and prices.

    Still deciding?

    The question that settles it is usually arithmetic rather than architecture: how many times will you read the whole corpus, and what does each of those reads cost where it currently lives. The cost of making a media library searchable walks through the four cost centres, and self-hosted Mixpeek covers running the same stack inside your own cloud account rather than on your own metal.