Live demo

7,141 labelled actions in. 2 out.

One agent on the "Act with approval" rung, 90 days, about eighty actions a day. Its true error rate is 1.2% against a 2.0% bar. On day 60 a model update takes it to 3.5%. Nobody typed in the day it was promoted or the day it was demoted; the test did.

Promoted
day 18
16 errors in 1,472 actions
Demoted
day 70
10 days after the update
Evidence to decide
3.00
ln(1/alpha) at alpha 0.05
Rung today
Act with approval
bar 2.0%
Evidence path, sampled every five days. Positive favours below the bar; the other column favours above it.
DayStatusActions in testErrorsRateBelow barAbove bar
4Act with approval · testing40771.7%-0.29-0.97
9Act with approval · testing79191.1%1.31-1.89
14Act with approval · testing1,211151.2%1.51-2.16
19Act unsupervised · watching12632.4%-0.44-0.21
24Act unsupervised · watching522132.5%-1.23-0.21
29Act unsupervised · watching923131.4%0.48-1.77
34Act unsupervised · watching1,321181.4%0.99-2.11
39Act unsupervised · watching1,702241.4%1.04-2.27
44Act unsupervised · watching2,115281.3%2.05-2.56
49Act unsupervised · watching2,501321.3%2.95-2.78
54Act unsupervised · watching2,900391.3%2.71-2.84
59Act unsupervised · watching3,266441.3%3.07-2.96
64Act unsupervised · re-qualifying415112.7%-1.21-0.02
69Act unsupervised · re-qualifying830263.1%-2.111.59
74Act with approval · demoted1,200403.3%-2.603.55
79Act with approval · demoted1,602563.5%-3.026.26
84Act with approval · demoted1,997693.5%-3.247.56
89Act with approval · demoted2,403793.3%-3.327.21
"7 clean days" promotes an over-budget agent
91.5%
Sequential test promotes it
0.0%

200 agents at a true rate of 3.0%, over the 2.0% bar, watched for 60 days. The rule of thumb reads 10 actions a day and promotes after 7 consecutive clean days. The sequential test reads every label and is bounded at 5.0% by construction.

DiagnosticsWhat the numbers depend on

Applied to every label instead of a sample, "7 clean days" promotes 3.5% of agents that are genuinely at 1.2%: the rule is uninformative at one volume and unreachable at the other. Across 200 agents at 1.2%, the sequential test promotes 92.5% within 60 days, median day 18; this scenario's day 18 is one draw from that.

The test assumes the agent is the same agent throughout, so evidence resets at the model update. Without the reset, the weeks of clean evidence before the update carry through: the same jump to 3.5% is not caught by day 89. A deployment you did not wire in is a drift you will find late.

With the reset, across 200 agents that jump to 3.5%, the fresh test demotes 76.0% within fourteen days, median day 8 after the deployment. This scenario's 10 days is one draw from that. Both sweeps assume outcomes are exchangeable within a deployment; an error rate that moves with the day of the week breaks that, and the guarantee with it.