Live demo
11 test steps in. 3 out.
Twenty runs of an unchanged suite. Three steps are genuinely nondeterministic and eight are stable — but at twenty runs, "has not flaked yet" and "will not flake" are not the same claim.
Pass rate today
61%
Replayed pinned
100%
Expected 1.6 CI attempts per green build, which is why the team added retries and stopped reading failures.
| Step | Distinct outputs | Flake rate | 95% upper | |
|---|---|---|---|---|
| load fixtures | 1 | 0.0% | 16.1% | |
| plan: choose tool | 3 | 10.0% | 30.1% | |
| call search | 1 | 0.0% | 16.1% | |
| rank results | 3 | 25.0% | 46.9% | |
| call order lookup | 1 | 0.0% | 16.1% | |
| compose reply | 3 | 10.0% | 30.1% | |
| guardrail check | 1 | 0.0% | 16.1% | |
| persist transcript | 1 | 0.0% | 16.1% | |
| emit metrics | 1 | 0.0% | 16.1% | |
| close ticket | 1 | 0.0% | 16.1% | |
| assert final state | 1 | 0.0% | 16.1% |
Note the upper bounds on the eight “stable” steps. None of them is zero, because twenty clean runs cannot prove a step never flakes. That gap is the difference between a measurement and a hope, and it is why the suite gets worse the more steps you add.