The alert

The whole product is this message

A fleet dashboard with forty lines on it is not a detector. It is a place to confirm something you already suspect. This arrives before anyone suspects anything, with the pool named and the healthy peers listed next to it.

S
SiliconAPP#eng-alerts
🔴 Thermal throttle events jumped +425%pool=a3-h100-04
Baseline
35.07
Now
183.94
Confidence
99.9%
Started
Aug 11
Fleet
inference-prod (40 pools)
Pool
a3-h100-04
Change
driver 550.90 → 555.42
Rolled
Aug 11, 03:10 UTC
Likely cause · The changepoint lands the night driver 555.42 reached pool a3-h100-04 and no other pool. Throttle events rose first, bandwidth utilisation and token rate followed within a day, and cost per million tokens rose because the pool bills the same hourly rate for less work. Every node in the pool still reports healthy and stayed in rotation the whole time, which is why liveness checks never fired.
  1. 1.Roll a3-h100-04 back to driver 550.90 and hold it for 48 hours.
  2. 2.Compare per-node throttle counts inside the pool to separate a fleet-level driver issue from two bad cards.
  3. 3.Check the rack inlet temperature log for Aug 11 before blaming the driver.
AcknowledgeFalse positiveCUSUM · q < 0.0001 · BH-corrected
S
SiliconAPP#eng-alerts
🔴 Cost per million tokens jumped +36.9%pool=a3-h100-04
Baseline
$2.84
Now
$3.88
Confidence
99.9%
Started
Aug 11
Fleet
inference-prod (40 pools)
Pool
a3-h100-04
Change
driver 550.90 → 555.42
Rolled
Aug 11, 03:10 UTC
Likely cause · The changepoint lands the night driver 555.42 reached pool a3-h100-04 and no other pool. Throttle events rose first, bandwidth utilisation and token rate followed within a day, and cost per million tokens rose because the pool bills the same hourly rate for less work. Every node in the pool still reports healthy and stayed in rotation the whole time, which is why liveness checks never fired.
  1. 1.Roll a3-h100-04 back to driver 550.90 and hold it for 48 hours.
  2. 2.Compare per-node throttle counts inside the pool to separate a fleet-level driver issue from two bad cards.
  3. 3.Check the rack inlet temperature log for Aug 11 before blaming the driver.
AcknowledgeFalse positiveCUSUM · q < 0.0001 · BH-corrected
Why the peers matter

Three other pools ran the identical SKU on the identical workload through the same window and did not move. That is what turns a number into an argument. Without the peer comparison, a 28% token-rate drop is a fleet-wide capacity question. With it, it is a driver rollback.