The alert

The whole product is this message

Everyone in this category ships a dashboard. Dashboards get checked when someone remembers to check them. This arrives whether or not anyone remembered, with the statistics already run and the deploy already correlated.

A
AgentSREAPP#eng-alerts
🔴 Tool-call error rate jumped +380%tier=enterprise
Baseline
2.06%
Now
9.90%
Confidence
99.9%
Started
Aug 13
Agent
support-triage
Deploy
v2.14.0
Shipped
Aug 12, 14:22 PT
Traces
312,880
Likely cause · The changepoint lands within 40 minutes of deploy v2.14.0, which rewrote the tool-selection preamble. Error and retry rates rose only on tier=enterprise, whose conversations carry a longer tool manifest — consistent with the model losing the correct tool in a longer list. Task volume is unchanged, so this would not appear on a throughput dashboard.
  1. 1.Diff the tool-selection preamble between v2.13.4 and v2.14.0.
  2. 2.Replay the 40 flagged enterprise traces against v2.13.4 to confirm.
  3. 3.If confirmed, roll back the preamble and keep the rest of v2.14.0.
AcknowledgeFalse positiveCUSUM · q < 0.0001 · BH-corrected
A
AgentSREAPP#eng-alerts
🔴 Retry rate jumped +217%tier=enterprise
Baseline
4.31%
Now
13.7%
Confidence
99.9%
Started
Aug 13
Agent
support-triage
Deploy
v2.14.0
Shipped
Aug 12, 14:22 PT
Traces
312,880
Likely cause · The changepoint lands within 40 minutes of deploy v2.14.0, which rewrote the tool-selection preamble. Error and retry rates rose only on tier=enterprise, whose conversations carry a longer tool manifest — consistent with the model losing the correct tool in a longer list. Task volume is unchanged, so this would not appear on a throughput dashboard.
  1. 1.Diff the tool-selection preamble between v2.13.4 and v2.14.0.
  2. 2.Replay the 40 flagged enterprise traces against v2.13.4 to confirm.
  3. 3.If confirmed, roll back the preamble and keep the rest of v2.14.0.
AcknowledgeFalse positiveCUSUM · q < 0.0001 · BH-corrected
The ordering rule

Statistics decide whether something happened. The model only explains what already cleared that bar, and only ever sees a confirmed changepoint. Feeding raw numbers to a model and asking “is this an anomaly?” is how these products end up confidently wrong.