A public agent leaderboard that corrects itself

The agent at the top of every leaderboard ran nine tasks

Public boards rank by raw success rate, which hands first place to whoever attempted fewest. The Reliability Index shrinks every entrant toward the field, publishes the method, and puts the newcomer where it belongs.

No spam. One email when it is ready to try.

The problem

Leaderboards with unequal sample sizes are ranking luck and everyone reads them anyway

An entrant with nine attempts and nine successes tops a board over one with eight hundred attempts at 82%. Everybody senses this is wrong and no board fixes it, because the fix looks like penalising newcomers. So procurement decisions get made on a ranking that is partly noise, and the vendors who know it stay quiet.

9
tasks attempted by the entrant topping the raw board
600+
attempted by the entrants it displaces
0
public agent boards applying a correction

The insight

Shrinkage is not a penalty. It is the correct amount of belief to place in nine attempts

Empirical Bayes fits a prior from the field itself and pulls every entrant toward it in proportion to how little evidence they have brought. An entrant with eight hundred tasks barely moves. One with nine moves almost all the way. This is a sixty-year-old result that beats raw estimates on total error every time, and it is absent from every public board in the industry.

Method

Beta-binomial empirical Bayes with the prior fitted from the field by moments, per-entrant shrinkage weighted by attempts, credible intervals reported alongside every rate, and rank movement published so the correction is visible rather than hidden.

How it works

Four steps, no data science team

01
Submit runs

Open submission, published task set, results reproducible from the logs.

02
Fit the field

The prior comes from the entrants themselves, refitted each edition.

03
Publish both

Raw and corrected, side by side, with the rank change shown.

04
Show the interval

A rate without an interval is the thing that caused the problem.

Who it is for

Free to read, free to enter. The audience is the asset

Nobody pays. Buyers reading it are procurement teams choosing an agent; entrants are vendors who want a credible number.

Pricing

Read
$0
The index, the methodology, the historical editions.
  • Full rankings
  • Published method
  • Reproducible from logs
Most common
Enter
$0
Submit your agent. Same correction as everyone else.
  • Open submission
  • Credible intervals
  • Rank-change transparency
Private board
$1,200/mo
The same correction on your internal agent comparisons.
  • Private leaderboards
  • Your own task set
  • API access
  • Historical tracking

Competition

What exists, and what it does not do

WhoWhat they doThe gap
Public agent benchmarksRank agents on standardised task suites.Every one of them ranks on raw rates with unequal attempt counts, which is the specific error this exists to correct.
Artificial Analysis and model comparison sitesTrack model speed, price and quality.Models rather than agents, and comparing on quality without correcting for sample size.
Vendor-published numbersEvery vendor quotes their own result.Incomparable by construction, and nobody publishes the attempt count next to the rate.
Ignoring leaderboardsWhat sophisticated buyers already do.Rational given the current state, and it leaves them with no comparison at all.
How this fails

This generates no revenue and consumes real time on a schedule — an index that skips an edition is dead. It is also a target: the first vendor unhappy with their corrected rank will attack the methodology publicly, which is survivable only if the method is genuinely defensible and published in full. That is the reason to do it, though. Being the person who runs the credible index is worth more to everything else in this portfolio than any subscription, and it is the only asset here that compounds while you sleep.

Market

Not a market. Distribution, and the reason anyone believes the paid products

Zero direct revenue by design. Its value is that procurement teams cite it, vendors argue about it, and both groups then know who built it.

Enter the index

No spam. One email when it is ready to try.

Or just go look at the demo first →