Every one of these exists because a decision is being made on a number that will not support it. Did the thing really change. Is the eval big enough to tell. Did the campaign cause anything. Is the leader actually leading.
Start here
The rest of this page is the full set. These are the ones worth ten minutes.
Decides what an agent is allowed to do, with a guaranteed error rate. The only one here that makes an agent safe to ship rather than easier to debug.
Fastest path to a real user. Finds the week your agent got worse, without labels, and stays quiet the rest of the time.
Free, no signup. Can your eval even see the regression you are worried about? Run the numbers before you argue about them.
The shared core
Nine methods, most of them decades old, none of them in the toolkit of the product engineers who own this data. Every product below runs on them client-side, with no backend.
statistics → decide whether something happened
↓
model → explain the thing that already cleared the bar
The model is never asked whether something is an anomaly.
Reversing that order is how these products hallucinate.Your agent got worse three weeks ago. Nothing alerted.
The bet. The engine ports almost verbatim, every design partner is within twenty miles, and you are your own first user.
Your eval set cannot see the regression you are looking for.
The bet. The free calculator is the distribution. Every engineer who has argued about a three-point delta is a user.
"The model got worse this week" is a feeling. Give it a date.
The bet. A public dashboard is repeatable front-page distribution for a founder nobody has heard of.
One prompt edit tripled one route. You find out on the invoice.
The bet. Ingests provider usage exports, so it works on day one with zero integration.
Your payments vendor changed last Monday. Nobody told you.
The bet. Maps onto a named YC request for startups, and it is the half of that request nobody is building.
Nobody can tell you what your escalation threshold buys.
The bet. The only one of the thirty that makes an agent safe to deploy rather than easier to debug. That is a different budget.
One pool got 28% slower. The fleet average moved 0.7%.
The bet. Every company training or serving its own models has this problem and no tool for it.
You found 104 bugs. Nobody in the room can tell you how many are left.
The bet. It answers the one question every security review ends on and none of them answer. Nobody else is even trying.
You checked the experiment every morning. That broke it.
The bet. The purest statistics play of the thirty, and the failure it fixes is close to universal.
Two channels, both at 37% churn. Nothing else about them is the same.
The bet. Survival analysis is standard in biostatistics and absent from every growth tool. That gap is the whole company.
Your trial is decided on day 3 or on day 13. The conversion rate cannot tell you which.
The bet. It tells a PLG team which day of the trial to intervene on. A conversion percentage never has.
Most of what the platform took credit for would have happened anyway.
The bet. Every CFO already suspects the attributed number is fiction. This is the first thing that measures it.
Before/after credited your rollout with a trend that hit every region.
The bet. Most decisions that matter cannot be randomised. This is the estimator for all of them.
Retention is flat. The curve is not.
The bet. Retention is the metric every board asks about and the one nobody can localise in time.
This ad stopped working on August 10. You killed it on the 22nd.
The bet. The leading-indicator gap is real money: engagement breaks a week before conversion does.
Traffic is down. You have two weeks of arguing ahead of you.
The bet. Answer engines are reshaping organic traffic right now, and nobody can date their own losses.
Your p99.9 is a six-week event, not a rare one.
The bet. Insurance has priced tails this way for a century. Infrastructure sizes capacity off the worst week it happens to remember.
Your 90% forecast interval covers 58.5% of the time.
The bet. It wraps whatever forecaster a team already runs. Nothing to replace is the shortest path to adoption.
Your best rep closed three deals.
The bet. Every company has a leaderboard it acts on. Almost every one of them is ranking noise at the top.
Your supplier is not late yet. They are already inconsistent.
The bet. Variance moves before the mean does, which is a genuine head start nobody is taking.
By the time it has a category name, it has been running two weeks.
The bet. An emerging issue is visible in the queue days before it has a category name.
Revenue grew. Blended margin moved half a point. You lost the segment.
The bet. Revenue can grow while the business quietly gets worse, and blended margin hides it.
The vendor says 35% deflected. Volume fell on every queue that month
The bet. Every AI support agent comes up for renewal, and the person who signed has one number and it came from the vendor.
Most of the queue was never going to be a filing. Someone still has to read all of it.
The bet. Highest contract values of the thirty, and it sits inside a named YC request for startups.
The payer changed the rules in July. Your aging report will say so in September.
The bet. Unglamorous, enormous, and the pain converts directly into cash the buyer can count.
The attack started three weeks before your chargeback rate moved.
The bet. Fires while the metric everyone watches is still flat, because chargebacks lag by weeks.
Every test passed. The number in the board deck was still wrong.
The bet. The incumbents proved the category is worth billions. The wedge is price and dbt-nativeness.
You sample forty transactions at random. There are twenty thousand.
The bet. It decides which forty of forty thousand transactions a human looks at. Same audit hours, better spent.
One site in forty is behaving differently. Which one?
The bet. Regulators explicitly encourage central statistical monitoring. Almost nobody sells it to mid-size sponsors.
The deduction stopped applying in August. Nobody looked until March.
The bet. Payroll errors are structural and repeat every cycle until someone notices. Nobody is watching between runs.
You find out about shrink in January. It started in August.
The bet. Shrink is measured a few times a year. The leading indicators are in daily data already being collected.
Your agent has a key that opens everything
The bet. The security review that precedes every agent launch is a forcing event nobody currently has an answer for.
Your agent has an attack surface and no regression test for it
The bet. Tool use is the new attack surface and nobody has a regression test for it.
You installed a tool server. Nobody read what it tells your model to do
The bet. A public registry cited in install decisions is the cheapest distribution available to an unknown founder.
Most of your codebase was written by a model and nobody recorded which parts
The bet. Every acquisition and every security questionnaire will eventually ask which lines a human read.
A third of your agent decisions cannot be explained after the fact
The bet. Enforcement timetables are published. The question arrives on a date somebody has already set.
Your agent pasted a key into a reply. The DLP rule flagged two hundred other things
The bet. Every agent that reads tool results will eventually paste one back out, and the regex queue is already ignored by the time it does.
"It has been clean for a week" is not evidence about a 2% error rate
The bet. Every team that lets an agent act has to decide when it may act alone, and today that decision is a feeling with a calendar attached.
Your agent test suite passes 61% of the time on unchanged code
The bet. Everyone building agents has flaky tests and the only current answer is a retry counter.
Everyone routes to the cheap model on a hunch
The bet. Inference is the fastest-growing line on the bill and every routing rule is a hunch.
Generated code fails differently, and your linter was built for the other kind
The bet. Review is the bottleneck now, and precision is the only thing that makes a reviewer read the output.
You wrote the spec in English and the agent said it was done
The bet. The spec is the primary artifact now and nothing verifies the code still matches it.
You are paying for 600,000 tokens to get the value of 24,000
The bet. Context is simultaneously the biggest cost lever and the one chosen entirely by guesswork.
Your 30-second timeout is a number somebody typed
The bet. Every agent in production has this number typed into a config, and every hung call pays for it.
Your canary is a t-test you run every morning
The bet. Every prompt and model upgrade ships behind a traffic split, and the decision at the end of it is a t-test somebody runs each morning.
Your error rate is 4%. Which segments are actually worse?
The bet. The group-by is the first thing every team opens when a customer complains, and every one of them is sorted by the wrong number.
You paid $4,000 an environment for three that teach the exploit
The bet. RL environments are one of the few genuinely large and urgent budgets, and nobody is checking the deliveries.
Your eval suite says 96% and cannot detect anything
The bet. Every vertical AI company builds the same saturated eval suite and none of them know it cannot detect anything.
Your benchmark went up two and a half points because the model had seen the answers
The bet. Eval credibility rests on trust today, and the first serious public dispute will end that.
The agent at the top of every leaderboard ran nine tasks
The bet. Running the credible public index is worth more to everything else than any subscription.
You are running experiments you were never taught to interpret
The bet. The gap this portfolio keeps finding is a teaching market before it is a tooling market.
Can your eval see the regression you are worried about?
The bet. The cheapest distribution available: a free tool used by exactly the people who buy the paid ones.
Your judge passes a third of the failures. Your pass rate is wrong.
The bet. Every team with an LLM judge reports a number that is wrong by an amount they have never measured, and the fix is one formula.
The aggregate passed. Three customer slices did not.
The bet. Every team with an eval gate has been burned by the aggregate at least once, and the fix is a short function nobody has shipped as a gate.
Your retrieval recall is 98% against a pool your retrievers built
The bet. Every RAG team has a recall number that never moves, and the engineer who owns it already suspects why.
Firing the bottom five annotators fires the wrong one
The bet. Every labelled dataset has a bottom-five rule and none of them have a test; the re-label order is a deliverable the vendor cannot write.
Your agent works in the demo and nobody can say whether it is safe to ship
The bet. Revenue in week two, a domain you actually know, and three customers writing the product spec instead of a guess.
You are about to wire on a retention chart nobody checked
The bet. Funds pay quickly, decide fast, and have nobody who checks the statistics behind the growth claims.
You have 62 test cases and you know that is not enough
The bet. The shortest path in this entire portfolio from a warm introduction to a paid invoice.
Your customer wants an accuracy guarantee and nobody can price one
The bet. Genuinely nobody is doing this, and the reason is that it is hard rather than that it is unwanted.
How to read this
It is to stop betting a year on one guess. 60 landing pages, 60 pitches, one shared engine, and a signup form on each. Which one people give you their email for is a better signal than which one felt most exciting to build.
Everything here is a prototype. Real math on synthetic data with a known injected ground truth, real positioning, real capture. No auth, no billing, no customer data. The one that gets traction is the one that earns a backend.