Six ideas that fix how you read your own numbers
You are running experiments you were never taught to interpret
Not a statistics course. Six specific mistakes that engineers shipping models make every week, each one demonstrated live in the browser on data where the right answer is known.
The problem
Every AI team now runs experiments and almost none of them have a statistician
Checking a running experiment every morning, comparing two eval scores with a t-test, scoring twelve slices separately, treating a confidence score as a probability. Each is a specific, named, well-understood error, and each is committed daily by people who are excellent engineers and were simply never taught this. The gap is not intelligence. It is that nobody ever showed them.
The insight
A demonstration on data with a known answer beats an explanation every time
Telling someone that peeking inflates the false-positive rate produces agreement and no behaviour change. Showing them two hundred experiments with no effect at all, a third of which their current process would call winners, changes what they do on Monday. Every lesson runs live on synthetic data where the truth is injected, so the reader watches their own method fail rather than being told it might.
Each lesson runs the real estimator in the browser against data with injected ground truth: sequential testing on A/A trials, power on paired evals, BH across slices, conformal on an uncalibrated score, shrinkage on a leaderboard, and survival on censored cohorts.
How it works
Four steps, no data science team
Short, specific, and about a mistake rather than a technique.
Every lesson is interactive. Change the numbers and watch the conclusion change.
The calculators are free and stand alone.
Replies come to a person, and the questions set the next lesson.
Who it is for
Free. The reader is the audience for everything else
Engineers shipping model-backed products who have argued about whether a number was real and had no way to settle it.
Pricing
- –Six lessons
- –Interactive demonstrations
- –No email required
- –Power and sample size
- –Sequential testing
- –Runs in your browser
- –Half day, your data
- –Recorded
- –Follow-up review
Competition
What exists, and what it does not do
| Who | What they do | The gap |
|---|---|---|
| University statistics courses | Teach the subject properly, from foundations. | Twelve weeks, no context about evals or agents, and nobody shipping a product will take one. |
| Experimentation vendor blogs | Statsig and Eppo write well about sequential testing. | Genuinely good and it is marketing for their platform, so it stops where their product stops. |
| Asking a model | What most engineers do now. | Gives a correct explanation and no demonstration, and the reader still cannot tell whether their own case is affected. |
| Learning it after the incident | The current curriculum. | Effective and expensive. |
It takes months to matter and there is no way to accelerate that — an audience compounds or it does not exist. It also competes for the same hours as the products, and the failure mode is a founder who writes six good lessons and ships nothing. The argument for doing it anyway is that the material is already built: every lesson is a demo that exists in this repository, so the marginal cost is the writing, and it is the only asset here that keeps working when attention moves elsewhere.
Market
No revenue by design, other than occasional workshops
The audience is the point. Readers become the distribution for Abstain, EvalPower and the Index, and it is the cheapest customer acquisition available to a founder nobody has heard of.