Busy is not the same as working.
Proof checks every training run and inference endpoint while it runs. Within minutes it tells you whether a slowdown is the job or the infrastructure, and shows the evidence.
You pay per GPU-hour. You get per productive hour.
Divide the list price by the share of time the GPU actually computes. That is the number that decides your budget, and the one Proof moves.
Dashboards show symptoms. We return verdicts.
You can assemble DCGM, Prometheus and Grafana yourself and still be staring at thirty charts at 3 a.m. The hard problem is correlation: joining GPU counters, kernel traces, collectives, scheduler state and the training log into one answer. The most expensive failures produce no hardware symptom at all.
One agent per node. One verdict per incident.
Because the join happens at collection time, correlation is a lookup, not a forensic project.
Why operators run Proof on their own fleet.
Not another tool to resell. Leverage for the operator itself: support, SLAs, fleet health, and what your GPU-hour is worth against everyone else's.