A public agent leaderboard that corrects itself
The agent at the top of every leaderboard ran nine tasks
Public boards rank by raw success rate, which hands first place to whoever attempted fewest. The Reliability Index shrinks every entrant toward the field, publishes the method, and puts the newcomer where it belongs.
The problem
Leaderboards with unequal sample sizes are ranking luck and everyone reads them anyway
An entrant with nine attempts and nine successes tops a board over one with eight hundred attempts at 82%. Everybody senses this is wrong and no board fixes it, because the fix looks like penalising newcomers. So procurement decisions get made on a ranking that is partly noise, and the vendors who know it stay quiet.
The insight
Shrinkage is not a penalty. It is the correct amount of belief to place in nine attempts
Empirical Bayes fits a prior from the field itself and pulls every entrant toward it in proportion to how little evidence they have brought. An entrant with eight hundred tasks barely moves. One with nine moves almost all the way. This is a sixty-year-old result that beats raw estimates on total error every time, and it is absent from every public board in the industry.
Beta-binomial empirical Bayes with the prior fitted from the field by moments, per-entrant shrinkage weighted by attempts, credible intervals reported alongside every rate, and rank movement published so the correction is visible rather than hidden.
How it works
Four steps, no data science team
Open submission, published task set, results reproducible from the logs.
The prior comes from the entrants themselves, refitted each edition.
Raw and corrected, side by side, with the rank change shown.
A rate without an interval is the thing that caused the problem.
Who it is for
Free to read, free to enter. The audience is the asset
Nobody pays. Buyers reading it are procurement teams choosing an agent; entrants are vendors who want a credible number.
Pricing
- –Full rankings
- –Published method
- –Reproducible from logs
- –Open submission
- –Credible intervals
- –Rank-change transparency
- –Private leaderboards
- –Your own task set
- –API access
- –Historical tracking
Competition
What exists, and what it does not do
| Who | What they do | The gap |
|---|---|---|
| Public agent benchmarks | Rank agents on standardised task suites. | Every one of them ranks on raw rates with unequal attempt counts, which is the specific error this exists to correct. |
| Artificial Analysis and model comparison sites | Track model speed, price and quality. | Models rather than agents, and comparing on quality without correcting for sample size. |
| Vendor-published numbers | Every vendor quotes their own result. | Incomparable by construction, and nobody publishes the attempt count next to the rate. |
| Ignoring leaderboards | What sophisticated buyers already do. | Rational given the current state, and it leaves them with no comparison at all. |
This generates no revenue and consumes real time on a schedule — an index that skips an edition is dead. It is also a target: the first vendor unhappy with their corrected rank will attack the methodology publicly, which is survivable only if the method is genuinely defensible and published in full. That is the reason to do it, though. Being the person who runs the credible index is worth more to everything else in this portfolio than any subscription, and it is the only asset here that compounds while you sleep.
Market
Not a market. Distribution, and the reason anyone believes the paid products
Zero direct revenue by design. Its value is that procurement teams cite it, vendors argue about it, and both groups then know who built it.