Live demo
7,141 labelled actions in. 2 out.
One agent on the "Act with approval" rung, 90 days, about eighty actions a day. Its true error rate is 1.2% against a 2.0% bar. On day 60 a model update takes it to 3.5%. Nobody typed in the day it was promoted or the day it was demoted; the test did.
| Day | Status | Actions in test | Errors | Rate | Below bar | Above bar |
|---|---|---|---|---|---|---|
| 4 | Act with approval · testing | 407 | 7 | 1.7% | -0.29 | -0.97 |
| 9 | Act with approval · testing | 791 | 9 | 1.1% | 1.31 | -1.89 |
| 14 | Act with approval · testing | 1,211 | 15 | 1.2% | 1.51 | -2.16 |
| 19 | Act unsupervised · watching | 126 | 3 | 2.4% | -0.44 | -0.21 |
| 24 | Act unsupervised · watching | 522 | 13 | 2.5% | -1.23 | -0.21 |
| 29 | Act unsupervised · watching | 923 | 13 | 1.4% | 0.48 | -1.77 |
| 34 | Act unsupervised · watching | 1,321 | 18 | 1.4% | 0.99 | -2.11 |
| 39 | Act unsupervised · watching | 1,702 | 24 | 1.4% | 1.04 | -2.27 |
| 44 | Act unsupervised · watching | 2,115 | 28 | 1.3% | 2.05 | -2.56 |
| 49 | Act unsupervised · watching | 2,501 | 32 | 1.3% | 2.95 | -2.78 |
| 54 | Act unsupervised · watching | 2,900 | 39 | 1.3% | 2.71 | -2.84 |
| 59 | Act unsupervised · watching | 3,266 | 44 | 1.3% | 3.07 | -2.96 |
| 64 | Act unsupervised · re-qualifying | 415 | 11 | 2.7% | -1.21 | -0.02 |
| 69 | Act unsupervised · re-qualifying | 830 | 26 | 3.1% | -2.11 | 1.59 |
| 74 | Act with approval · demoted | 1,200 | 40 | 3.3% | -2.60 | 3.55 |
| 79 | Act with approval · demoted | 1,602 | 56 | 3.5% | -3.02 | 6.26 |
| 84 | Act with approval · demoted | 1,997 | 69 | 3.5% | -3.24 | 7.56 |
| 89 | Act with approval · demoted | 2,403 | 79 | 3.3% | -3.32 | 7.21 |
200 agents at a true rate of 3.0%, over the 2.0% bar, watched for 60 days. The rule of thumb reads 10 actions a day and promotes after 7 consecutive clean days. The sequential test reads every label and is bounded at 5.0% by construction.
Applied to every label instead of a sample, "7 clean days" promotes 3.5% of agents that are genuinely at 1.2%: the rule is uninformative at one volume and unreachable at the other. Across 200 agents at 1.2%, the sequential test promotes 92.5% within 60 days, median day 18; this scenario's day 18 is one draw from that.
The test assumes the agent is the same agent throughout, so evidence resets at the model update. Without the reset, the weeks of clean evidence before the update carry through: the same jump to 3.5% is not caught by day 89. A deployment you did not wire in is a drift you will find late.
With the reset, across 200 agents that jump to 3.5%, the fresh test demotes 76.0% within fourteen days, median day 8 after the deployment. This scenario's 10 days is one draw from that. Both sweeps assume outcomes are exchangeable within a deployment; an error rate that moves with the day of the week breaks that, and the guarantee with it.