FAIL or a WARN, find the row here.
Substrate
Rows fromsubstrate_e2e.py, which runs before anything of ours is installed.
pvc-durability: the volume was empty on a second pod
Writes do not survive the pod that made them. MVS keeps its index on a
PersistentVolume, so this is disqualifying rather than inconvenient. Usually the
claim bound to something ephemeral, or the StorageClass provisioner is not
actually provisioning.
pvc-durability: the volume returned a marker that is not this run’s
The claim came back carrying data the suite did not write. Durability was not
demonstrated, and separately, a volume being reused across workloads is worth
understanding before anything of ours stores an index on it. Check whether the
StorageClass is handing out a pre-existing volume.
pvc-durability: SKIP, this identity cannot create PVCs
The row needs to create one claim and two pods. Re-run as an identity with
create on persistentvolumeclaims in the namespace, or accept that the
durability claim is untested.
storage-class: no StorageClass exists
Nothing can provision a PersistentVolume, so MVS cannot run on this cluster as it
stands.
storage-class: WARN, no class is marked default
A PVC without an explicit storageClassName will never bind. Tell us which class
to use and we render it into the chart rather than assuming.
storage-class: WARN, allowVolumeExpansion is false
An MVS shard that outgrows its volume has to be replaced rather than expanded.
Workable, and worth knowing before it happens at 3am.
storage-throughput: WARN, no floor given
The row measured a number and asserted nothing about it, which is what it does
without --require-write-mbps. Ask us for the floor your configuration needs.
storage-throughput: FAIL, below the stated floor
Worth reading the caveat with the number: this is a sequential write to a freshly
provisioned volume, measured once. Cloud block storage sells burst credits, so
sustained indexing will be slower than the figure that already failed. If the
class is sized by capacity (GKE pd-balanced, gp3 without provisioned IOPS), a
larger volume is often the whole fix.
workload-identity: WARN, no ServiceAccount is bound to a cloud identity
Mixpeek would need static credentials in a Secret to reach object storage. That
is workable and is a posture worth choosing on purpose rather than by default. If
you have wired an identity some other way, this row cannot see it: it reads
ServiceAccount annotations only and does not attempt a bucket call. The Ray
suite’s object-storage row does that end to end, after install.
crd-install-grant: WARN, this identity cannot create ClusterRoles
Not a failure, and both answers are supported. It means your platform team
installs and upgrades KubeRay, and the cost is that a Mixpeek release requiring a
newer KubeRay needs a change on your side first. See
Security and RBAC.
On GKE, note that container.customResourceDefinitions.create and
container.clusterRoles.create are separate permissions. Holding one does not
imply the other, which is why the row checks three verbs and not one.
namespace-grant: FAIL, cannot create some resource in the namespace
The install creates it, so an install would fail partway through and leave
objects behind. The row names each missing verb.
pod-security: WARN, the namespace enforces restricted
Our pods must run non-root with a seccompProfile and no added capabilities.
Tell us before the install so the chart renders a compliant securityContext.
Otherwise the first failure is an admission webhook message that reads like a
Mixpeek bug.
node-capacity: FAIL, not enough untainted allocatable capacity
Tainted and cordoned nodes are excluded and listed in the evidence. If Mixpeek
should be scheduling onto them, tell us the tolerations and that capacity counts.
A cluster that is merely fenced should not read as too small.
node-capacity: SKIP, no requirement given
The row will not infer its own bound. Ask us for the CPU and memory your
configuration needs and pass them.
kuberay: WARN, not installed
Expected on a cluster where KubeRay is not installed. Somebody has to install
it; crd-install-grant says who can.
Cluster
ray-reachable: UNREACHABLE, or the dashboard does not answer
The suite could not reach the Ray dashboard at the address it was given. Nothing
downstream ran, so the rest of the report is empty rather than clean.
Usually one of: the dashboard is not exposed outside the cluster, a port-forward
died, or the address points at a different cluster. Confirm with
curl <dashboard>/api/version from wherever the suite runs.
fingerprint: the cluster is not the one you think
The suite ran a job on the target and compared what came back against the address
it was given. A mismatch means two different clusters are involved, which makes
every other row describe something you did not intend to test.
kuberay: CRDs not served, or below 1.1
KubeRay 1.1 or newer is required. Below it, submissionMode: HTTPMode does
not exist and batch cannot be submitted at all. Check with
kubectl get crd rayjobs.ray.io.
Security
rbac-minimal-sa: a forbidden verb answered yes
The submitter ServiceAccount can do something it must not: create a
ClusterRole, create namespaces, delete nodes, or create RayJobs outside its own
namespace. The row names which.
This usually means the kit was applied on top of a pre-existing binding that
granted more, or the ServiceAccount was reused. It is the one row where a FAIL
means you have given us too much rather than too little.
rbac-minimal-sa: every answer is no
kubectl auth can-i --as needs impersonation rights. Run as an identity that
cannot impersonate, everything answers no, and the row reads like a clean bill
of health. The suite calls this out; re-run as an identity that can impersonate.
rayjob-cr-secrets: FAIL
Expected today, and it is our defect rather than yours. Batch RayJob CRs carry
credential-shaped environment variables inline, so anyone with get/list on
rayjobs can read them. See Security and RBAC for the
fix in flight and the two-minute TTL that bounds it. The row stays red until it
lands, deliberately.
network-policy: WARN, no NetworkPolicy exists
Not a failure. Mixpeek’s workloads in that namespace have unrestricted egress,
which is a legitimate choice on a flat internal cluster. The row exists so that
choosing it and forgetting stop looking the same.
network-policy: WARN, every policy carries a podSelector
A NetworkPolicy exists but the suite cannot resolve whether it selects our
pods. A policy that selects something else leaves our workloads as open as none at
all while still counting as one. Confirm the selector matches the pods in the
Mixpeek namespace.
Connectivity
registry-pull: the nodes could not pull the image
Your nodes, not the machine running the suite. Almost always a missing
imagePullSecret in the namespace, or a registry your nodes cannot authenticate
to. Mirror the image into a registry they already trust, or add the pull secret.
registry-pull / scheduling / failure-recovery: SKIP, cannot create pods
Not a defect in your cluster. The identity running the suite lacks create on
pods, which the submitter Role does not grant by design. Re-run those rows as an
identity that can create pods in the namespace.
object-storage: WARN, static keys in the environment
The bucket was reachable, but the job carried static credentials, so the run does
not demonstrate the cluster’s own workload identity. If you intend to use
Workload Identity, this is the row that tells you it is not actually in use yet.
control-plane: probe never reported from inside the cluster
The job that was supposed to reach your control-plane endpoints never came back.
Usually egress: a NetworkPolicy, a firewall rule, or a private endpoint without a
route. The suite runs this from inside the cluster on purpose, because reachable
from your laptop and reachable from a Ray worker are different questions.
engine-on-target: the engine claims services not on this dashboard
The Mixpeek engine you reached is running somewhere other than the cluster under
test. The report would otherwise describe two machines at once.
Workloads
scheduling: a pod without a toleration still landed
Your taint is not being honoured. The suite submitted the same pod twice,
differing only by a toleration, and both scheduled.
scheduling: did not schedule, but the reason does not mention a taint
The pod was blocked by something else, most often insufficient CPU. The row fails
because a pod failing to schedule is not evidence your taint did it, and crediting
it to the taint would report a working policy that was never exercised.
scheduling: a pod requesting 999 CPU was accepted
There is no effective ResourceQuota in the namespace. Nothing bounds what
Mixpeek can ask for.
autoscaling: SKIP, no maximum given
The row needs to be told your maxReplicas. It will not infer it: an assertion
that the count never exceeded the highest value observed is true by construction
and cannot fail.
autoscaling: went past maxReplicas
The autoscaler exceeded the ceiling in your own configuration. Worth escalating;
this is your policy not being honoured by Ray, not by Mixpeek.
autoscaling: scaled up and never came back down
Workers remained after the load finished. You are paying for idle capacity. Check
idleTimeoutSeconds in the cluster’s autoscaler options.
failure-recovery: two faults are indistinguishable
Two different infrastructure failures produced the same diagnosis through the API,
so a customer hitting one cannot tell which they hit. The row names the pair.
failure-recovery: a fault never reached a settled state
The suite read a transition rather than an outcome, and says so rather than
reporting the transition as a diagnosis.
serve-lifecycle: the customer’s app moved
Deploying, scaling or removing a Mixpeek Serve application changed one of yours:
its status, its replica count, or its deploy timestamp. This is the row that must
never fail on a shared cluster. See
Existing RayCluster for the deployment path and
why it is per-application.
serve-lifecycle: INCONCLUSIVE, cluster unstable
The suite watched your cluster before writing anything and something moved on its
own. Nothing can be attributed to Mixpeek in either direction on that run. Usually
memory pressure or a cluster still settling; re-run once it is quiet.
Isolation
non-interference: a foreign app, job or actor changed
Something of yours is no longer as it was after Mixpeek ran. The row names it. If
the cluster was also unstable, the row says INCONCLUSIVE instead, because
interference and a cluster falling over produce the same observation.
cleanup: Mixpeek residue after uninstall
Objects we created are still present after teardown. The row lists them.
Verdicts
INCOMPLETE
Nothing failed and not everything ran. Look for SKIP rows: a missing credential,
a permission the running identity lacks, or a check that could not run. It is not a
pass and not a failure, and treating it as either is the mistake it exists to
prevent.

