The alert
The whole product is this message
Everyone in this category ships a dashboard. Dashboards get checked when someone remembers to check them. This arrives whether or not anyone remembered, with the statistics already run and the deploy already correlated.
A
AgentSREAPP#eng-alerts🔴 Tool-call error rate jumped +380%tier=enterprise
Baseline
2.06%
Now
9.90%
Confidence
99.9%
Started
Aug 13
Agent
support-triage
Deploy
v2.14.0
Shipped
Aug 12, 14:22 PT
Traces
312,880
Likely cause · The changepoint lands within 40 minutes of deploy v2.14.0, which rewrote the tool-selection preamble. Error and retry rates rose only on tier=enterprise, whose conversations carry a longer tool manifest — consistent with the model losing the correct tool in a longer list. Task volume is unchanged, so this would not appear on a throughput dashboard.
- 1.Diff the tool-selection preamble between v2.13.4 and v2.14.0.
- 2.Replay the 40 flagged enterprise traces against v2.13.4 to confirm.
- 3.If confirmed, roll back the preamble and keep the rest of v2.14.0.
AcknowledgeFalse positiveCUSUM · q < 0.0001 · BH-corrected
A
AgentSREAPP#eng-alerts🔴 Retry rate jumped +217%tier=enterprise
Baseline
4.31%
Now
13.7%
Confidence
99.9%
Started
Aug 13
Agent
support-triage
Deploy
v2.14.0
Shipped
Aug 12, 14:22 PT
Traces
312,880
Likely cause · The changepoint lands within 40 minutes of deploy v2.14.0, which rewrote the tool-selection preamble. Error and retry rates rose only on tier=enterprise, whose conversations carry a longer tool manifest — consistent with the model losing the correct tool in a longer list. Task volume is unchanged, so this would not appear on a throughput dashboard.
- 1.Diff the tool-selection preamble between v2.13.4 and v2.14.0.
- 2.Replay the 40 flagged enterprise traces against v2.13.4 to confirm.
- 3.If confirmed, roll back the preamble and keep the rest of v2.14.0.
AcknowledgeFalse positiveCUSUM · q < 0.0001 · BH-corrected
The ordering rule
Statistics decide whether something happened. The model only explains what already cleared that bar, and only ever sees a confirmed changepoint. Feeding raw numbers to a model and asking “is this an anomaly?” is how these products end up confidently wrong.