60 prototypes · two engines · 8 markets

60 products.
One statistician.

Every one of these exists because a decision is being made on a number that will not support it. Did the thing really change. Is the eval big enough to tell. Did the campaign cause anything. Is the leader actually leading.

No spam. One email when it is ready to try.

CUSUM · BOCDdid this actually change
Benjamini-Hochbergdid you test too many things
Bootstrap · mSPRTcould you even have seen it
Conformalhow sure are you, with a guarantee
DiD · Synthetic controldid you cause it
Kaplan-Meierwhen, not how many
Extreme value theoryhow bad can it get
Empirical Bayesis that rank real
Benford · Capturewhat is hidden, what did you miss

Start here

Three of the 60, if you only look at three.

The rest of this page is the full set. These are the ones worth ten minutes.

01Abstain
Nobody can tell you what your escalation threshold buys.

Decides what an agent is allowed to do, with a guaranteed error rate. The only one here that makes an agent safe to ship rather than easier to debug.

5.9% → 2.2%
error rate under a hand-picked threshold, then a calibrated one
02AgentSRE
Your agent got worse three weeks ago. Nothing alerted.

Fastest path to a real user. Finds the week your agent got worse, without labels, and stays quiet the rest of the time.

88 → 5
naive alerts vs statistical alerts
03The Calculators
Can your eval see the regression you are worried about?

Free, no signup. Can your eval even see the regression you are worried about? Run the numbers before you argue about them.

11.6pp → 2.4pp
smallest visible regression, 40 cases versus 900

The shared core

One engine, 60 front doors.

Nine methods, most of them decades old, none of them in the toolkit of the product engineers who own this data. Every product below runs on them client-side, with no backend.

01CUSUM · BOCDdid this actually change
02Benjamini-Hochbergdid you test too many things
03Bootstrap · mSPRTcould you even have seen it
04Conformalhow sure are you, with a guarantee
05DiD · Synthetic controldid you cause it
06Kaplan-Meierwhen, not how many
07Extreme value theoryhow bad can it get
08Empirical Bayesis that rank real
09Benford · Capturewhat is hidden, what did you miss
··Parity tests8 series · 5 BH sets · TS matches Python to 1e-6
the rule that runs through all of them
statistics  →  decide whether something happened
      ↓
    model     →  explain the thing that already cleared the bar

The model is never asked whether something is an anomaly.
Reversing that order is how these products hallucinate.

Dev tools & AI infrastructure

8 products
01AgentSREDrift detection for agents in production.

Your agent got worse three weeks ago. Nothing alerted.

88 → 5
naive alerts vs statistical alerts
14 series · 43 of the naive alerts fired before anything was wrong · 0 false positives

The bet. The engine ports almost verbatim, every design partner is within twenty miles, and you are your own first user.

02EvalPowerStatistical power for LLM evals.

Your eval set cannot see the regression you are looking for.

0.49 → 0.006
p-value on the same regression, at n=40 then n=600
A real 3-point regression that a 40-case eval reads as 5 points and cannot confirm, resolved cleanly at 600

The bet. The free calculator is the distribution. Every engineer who has argued about a three-point delta is a user.

03Drift RadarEvidence for when a model quietly changed.

"The model got worse this week" is a feeling. Give it a date.

44 → 5
naive alerts vs statistical alerts
14 series · 28 of the naive alerts fired before anything was wrong · 0 false positives

The bet. A public dashboard is repeatable front-page distribution for a founder nobody has heard of.

04SpendGuardCatch inference cost regressions in a day.

One prompt edit tripled one route. You find out on the invoice.

23 → 4
naive alerts vs statistical alerts
14 series · 12 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Ingests provider usage exports, so it works on day one with zero integration.

05Contract DriftKnow when a vendor API changed. Before support does.

Your payments vendor changed last Monday. Nobody told you.

38 → 4
naive alerts vs statistical alerts
14 series · 14 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Maps onto a named YC request for startups, and it is the half of that request nobody is building.

06AbstainCalibrated escalation for AI agents.

Nobody can tell you what your escalation threshold buys.

5.9% → 2.2%
error rate under a hand-picked threshold, then a calibrated one
A 0.9 confidence cutoff escalates 12.8% and delivers 5.9% errors against a stated 2% budget · conformal hits the budget at 34.6% escalation

The bet. The only one of the thirty that makes an agent safe to deploy rather than easier to debug. That is a different budget.

07SiliconFind the node group that got slower.

One pool got 28% slower. The fleet average moved 0.7%.

50 → 4
naive alerts vs statistical alerts
14 series · 28 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Every company training or serving its own models has this problem and no tool for it.

08UnknownEstimate what your testing did not find.

You found 104 bugs. Nobody in the room can tell you how many are left.

104 → 175
defects found, versus defects estimated to exist
71 still undiscovered at 59.3% coverage · estimated from method overlap, not from effort

The bet. It answers the one question every security review ends on and none of them answer. Nobody else is even trying.

Growth & revenue

8 products
09Experiment GuardA/B tests that survive being watched.

You checked the experiment every morning. That broke it.

36% → 1%
false-win rate under continuous monitoring
200 A/A experiments with zero true effect · 72 naive false wins vs 2 sequential

The bet. The purest statistics play of the thirty, and the failure it fixes is close to universal.

10Churn ClockSurvival analysis for subscription churn.

Two channels, both at 37% churn. Nothing else about them is the same.

0.90 → <1e-9
p-value on the same two cohorts, by percentage then by log-rank
Churn rates of 37.4% and 37.8% look identical · the cohorts churn at completely different times, and the percentage cannot see it

The bet. Survival analysis is standard in biostatistics and absent from every growth tool. That gap is the whole company.

11Trial PhysicsTrial conversion as a hazard curve.

Your trial is decided on day 3 or on day 13. The conversion rate cannot tell you which.

0.10 → 2.2e-11
p-value comparing two onboarding flows, by percentage then by hazard
The percentage comparison cannot call it while trials are still running · the survival test resolves it decisively on the same data

The bet. It tells a PLG team which day of the trial to intervene on. A conversion percentage never has.

12LiftGeo holdout incrementality for ad spend.

Most of what the platform took credit for would have happened anyway.

$2.15M → $448k
platform-attributed revenue versus measured incremental revenue
True injected lift was $490k · synthetic control recovers it, the platform overstates by 4.8x

The bet. Every CFO already suspects the attributed number is fiction. This is the first thing that measures it.

13RolloutCausal readouts for changes you cannot A/B test.

Before/after credited your rollout with a trend that hit every region.

1.02 → 0.42
naive before/after versus the causal estimate
A seasonal trend of 0.60 hit every region · the naive read credits all of it to the rollout, and difference-in-differences strips it out

The bet. Most decisions that matter cannot be randomised. This is the estimator for all of them.

14Retention CliffFind where the retention curve broke.

Retention is flat. The curve is not.

60 → 4
naive alerts vs statistical alerts
16 series · 32 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Retention is the metric every board asks about and the one nobody can localise in time.

15CreativeDated creative fatigue, not guessed.

This ad stopped working on August 10. You killed it on the 22nd.

27 → 4
naive alerts vs statistical alerts
13 series · 11 of the naive alerts fired before anything was wrong · 0 false positives

The bet. The leading-indicator gap is real money: engagement breaks a week before conversion does.

16OrganicDate the break in your organic traffic.

Traffic is down. You have two weeks of arguing ahead of you.

59 → 4
naive alerts vs statistical alerts
15 series · 31 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Answer engines are reshaping organic traffic right now, and nobody can date their own losses.

Operations & reliability

7 products
17TailCapacity planning with extreme value theory.

Your p99.9 is a six-week event, not a rare one.

1017ms → 2944ms
sample p99.9 versus the modelled once-a-year hour
Fitted shape 0.41 against a true 0.45 · the yearly event exceeds the worst hour in 90 days (2057ms), which a percentile can never tell you

The bet. Insurance has priced tails this way for a century. Infrastructure sizes capacity off the worst week it happens to remember.

18CoverageForecast intervals that are actually right.

Your 90% forecast interval covers 58.5% of the time.

58.5% → 90.7%
actual coverage of a nominal 90% interval, before and after
The model’s own 90% band covers barely half the time · conformal delivers what it promises, on the same forecaster, without touching it

The bet. It wraps whatever forecaster a team already runs. Nothing to replace is the shortest path to adoption.

19RankedLeaderboards that survive their own sample sizes.

Your best rep closed three deals.

0.311 → 0.041
total squared error against the true rates, raw then shrunk
The rep who tops the raw board closed 3 of 5 and is genuinely below average · shrinkage moves them to #11 and cuts total error 7.6x

The bet. Every company has a leaderboard it acts on. Almost every one of them is ranking noise at the top.

20Lead TimeSupplier drift, caught before the line stops.

Your supplier is not late yet. They are already inconsistent.

66 → 4
naive alerts vs statistical alerts
16 series · 33 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Variance moves before the mean does, which is a genuine head start nobody is taking.

21SupportSee the new issue before it has a tag.

By the time it has a category name, it has been running two weeks.

87 → 4
naive alerts vs statistical alerts
14 series · 46 of the naive alerts fired before anything was wrong · 0 false positives

The bet. An emerging issue is visible in the queue days before it has a category name.

22MarginCatch unit economics drifting by segment.

Revenue grew. Blended margin moved half a point. You lost the segment.

52 → 4
naive alerts vs statistical alerts
14 series · 24 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Revenue can grow while the business quietly gets worse, and blended margin hides it.

23DeflectTickets the support bot actually removed

The vendor says 35% deflected. Volume fell on every queue that month

35.0% → 17.4%
deflection the vendor dashboard claims, versus tickets that actually went away
159 tickets a week removed against 915 expected · before/after says 25.0% and the season is the difference · pre-launch tracking error 1.6%

The bet. Every AI support agent comes up for renewal, and the person who signed has one number and it came from the vendor.

Vertical AI & compliance

8 products
24AML TriageFDR control for the AML alert queue.

Most of the queue was never going to be a filing. Someone still has to read all of it.

73 → 4
naive alerts vs statistical alerts
14 series · 39 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Highest contract values of the thirty, and it sits inside a named YC request for startups.

25Denial RadarPayer policy changes, caught in days.

The payer changed the rules in July. Your aging report will say so in September.

89 → 4
naive alerts vs statistical alerts
14 series · 47 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Unglamorous, enormous, and the pain converts directly into cash the buyer can count.

26Fraud DriftCatch new attack patterns before the chargebacks.

The attack started three weeks before your chargeback rate moved.

41 → 4
naive alerts vs statistical alerts
15 series · 21 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Fires while the metric everyone watches is still flat, because chargebacks lag by weeks.

27Pipeline SentinelColumn-level data monitoring for dbt teams.

Every test passed. The number in the board deck was still wrong.

50 → 4
naive alerts vs statistical alerts
15 series · 30 of the naive alerts fired before anything was wrong · 0 false positives

The bet. The incumbents proved the category is worth billions. The wedge is price and dbt-nativeness.

28LedgerAnomaly screening for expense and AP data.

You sample forty transactions at random. There are twenty thousand.

19,800 → 3
transactions screened, versus units queued for human review
16 cost centres tested with FDR control · a digit deviation is a reason to look, never a finding of fraud

The bet. It decides which forty of forty thousand transactions a human looks at. Same audit hours, better spent.

29SitewatchFind the trial site that stopped looking normal.

One site in forty is behaving differently. Which one?

143 → 4
naive alerts vs statistical alerts
16 series · 77 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Regulators explicitly encourage central statistical monitoring. Almost nobody sells it to mid-size sponsors.

30PayrollCatch payroll errors before the next cycle.

The deduction stopped applying in August. Nobody looked until March.

25 → 4
naive alerts vs statistical alerts
13 series · 15 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Payroll errors are structural and repeat every cycle until someone notices. Nobody is watching between runs.

31ShrinkShrink localised to a store, a category, a week.

You find out about shrink in January. It started in August.

132 → 4
naive alerts vs statistical alerts
16 series · 69 of the naive alerts fired before anything was wrong · 0 false positives

The bet. Shrink is measured a few times a year. The leading indicators are in daily data already being collected.

Agent security & governance

7 products
32KeyringScoped, expiring credentials for agents

Your agent has a key that opens everything

156 → 31
blast radius granted, versus what tasks actually use
80.1% of granted blast radius never exercised in 2,400 tasks · all 3 task types now safe to tighten

The bet. The security review that precedes every agent launch is a forcing event nobody currently has an answer for.

33Injection RangePrompt-injection regression tests in CI

Your agent has an attack surface and no regression test for it

14 → 4
flagged by an uncorrected screen, versus after correction
64 tool descriptions with 4 planted payloads · every legitimate tool mentioning credentials or webhooks suppressed

The bet. Tool use is the new attack surface and nobody has a regression test for it.

34MCP AuditSecurity grades for the tool servers you install

You installed a tool server. Nobody read what it tells your model to do

12 → 4
tool servers graded, versus servers that fail
64 tool descriptions inspected · a malicious description needs no CVE, just a sentence

The bet. A public registry cited in install decisions is the cheapest distribution available to an unknown founder.

35ProvenanceLine-level attribution for AI-written code

Most of your codebase was written by a model and nobody recorded which parts

16,695 → 12,407
model-written lines, and the ones nobody reviewed
Mix shifted on 2026-07-27 at confidence 1.000 · review coverage 25.7% and flat throughout

The bet. Every acquisition and every security questionnaire will eventually ask which lines a human read.

36JustifyAn audit trail an examiner will accept

A third of your agent decisions cannot be explained after the fact

400 → 210
agent actions audited, versus actions that cannot be explained
107 of them carried a real consequence · completeness is a conjunction, so three elements of four is not partial credit

The bet. Enforcement timetables are published. The question arrives on a date somebody has already set.

37LeakCatch the six outputs that leaked, not 200

Your agent pasted a key into a reply. The DLP rule flagged two hundred other things

218 → 6
outputs flagged by regex rules, versus after calibration and correction
500 outputs with 6 planted leaks · 50 awkward legitimate outputs suppressed by correction · calibrated on 400 prior outputs

The bet. Every agent that reads tool results will eventually paste one back out, and the regex queue is already ignored by the time it does.

38LadderEarn agent autonomy with evidence, not a week

"It has been clean for a week" is not evidence about a 2% error rate

91.5% → 0.0%
over-budget agents promoted to unsupervised, under "seven clean days" versus the sequential test
Promoted day 18 at a true rate of 1.2% under a 2.0% bar · demoted day 70, 10 days after the model update on day 60

The bet. Every team that lets an agent act has to decide when it may act alone, and today that decision is a feeling with a calendar attached.

AI engineering practice

8 products
39ReplayDeterministic replay for agent tests

Your agent test suite passes 61% of the time on unchanged code

60.8% → 100%
suite pass rate on unchanged code, live then replayed
3 of 11 steps genuinely nondeterministic · 1.6 expected CI attempts per green build

The bet. Everyone building agents has flaky tests and the only current answer is a retry counter.

40RouterCost routing with a stated error rate

Everyone routes to the cheap model on a hunch

$5.00 → $1.10
cost per thousand requests, premium model versus routed
100.0% of routed traffic clears the quality bar against a promised 90.0% · the guarantee held on traffic the policy never saw

The bet. Inference is the fastest-growing line on the bill and every routing rule is a hunch.

41Second PassReview tuned to how generated code fails

Generated code fails differently, and your linter was built for the other kind

4 → 4
defects planted in the pull request, versus findings reported
34 lines across 4 files · 2 blocking, 0 style comments, no false positives

The bet. Review is the bottleneck now, and precision is the only thing that makes a reviewer read the output.

42Spec LockTurns a written spec into failing tests

You wrote the spec in English and the agent said it was done

4 of 5
spec lines that became executable properties
15 violations across 120 runs · 1 line reported as unfalsifiable rather than silently passed

The bet. The spec is the primary artifact now and nothing verifies the code still matches it.

43Context BudgetMeasures what is worth putting in context

You are paying for 600,000 tokens to get the value of 24,000

358k → 12.6k
tokens sent, everything versus what earns its place
98.1% of measured contribution retained · 12 chunks contribute nothing measurable at all

The bet. Context is simultaneously the biggest cost lever and the one chosen entirely by guesswork.

44TimeoutTool-call timeouts set by survival analysis

Your 30-second timeout is a number somebody typed

30s → 6s
timeout in the config, versus where waiting stops paying
5,000 calls, 15.8% still running at 30s · 0.4% of completions cut · 24s saved per hung call

The bet. Every agent in production has this number typed into a config, and every hung call pays for it.

45CanaryAlways-valid canaries for model upgrades

Your canary is a t-test you run every morning

22.0% → 3.0%
good canaries halted by mistake with daily peeking, versus sequential
200 A/A canaries, 14 daily looks each · a 3-point regression rolled back after 30,000 requests · guarantee holds at every look

The bet. Every prompt and model upgrade ships behind a traffic split, and the decision at the end of it is a t-test somebody runs each morning.

46UnevenWhich segments are really worse

Your error rate is 4%. Which segments are actually worse?

5 → 2
segments flagged by raw rate, versus real after shrinkage and correction
Population rate 4.5% over 23,240 requests · 3 of the 3 false flags had 60 requests or fewer · prior worth 64 requests

The bet. The group-by is the first thing every team opens when a customer complains, and every one of them is sorted by the wrong number.

Evaluation & measurement

10 products
47Environment ForgeVerification for RL task environments

You paid $4,000 an environment for three that teach the exploit

12 → 3
environments delivered, versus environments that are gameable
840 episodes · every rejection names the exploiting policy and requires two independent tells before firing

The bet. RL environments are one of the few genuinely large and urgent budgets, and nobody is checking the deliveries.

48Eval KitsDomain eval suites that can actually see a regression

Your eval suite says 96% and cannot detect anything

2.0 → 8.0
eval scorecard out of ten, suite today versus rebuilt
95.7% pass rate means variance is near zero, so the suite cannot see a regression of any size · 2 slices too thin to report

The bet. Every vertical AI company builds the same saturated eval suite and none of them know it cannot detect anything.

49ContaminationTests whether your benchmark leaked

Your benchmark went up two and a half points because the model had seen the answers

73.3% → 70.2%
benchmark score reported, versus on clean items only
22 of 180 items look memorised · permutation p < 0.001 over 4,000 relabellings

The bet. Eval credibility rests on trust today, and the first serious public dispute will end that.

50Reliability IndexA public agent leaderboard that corrects itself

The agent at the top of every leaderboard ran nine tasks

#1 → #3
where a nine-task entrant lands, raw board then corrected
13 entrants, 8,020 tasks · prior fitted from the field is worth 45 tasks, so established entrants barely move

The bet. Running the credible public index is worth more to everything else than any subscription.

51Statistics for AI EngineersSix ideas that fix how you read your own numbers

You are running experiments you were never taught to interpret

36.0% → 1.0%
false-win rate, checked daily versus always-valid
200 experiments with a byte-identical treatment · every win false by construction, and the reader watches their own method produce them

The bet. The gap this portfolio keeps finding is a teaching market before it is a tooling market.

52The CalculatorsFree tools for questions you keep guessing at

Can your eval see the regression you are worried about?

11.6pp → 2.4pp
smallest visible regression, 40 cases versus 900
A 40-case eval prints a number every time you run it and cannot see a 3-point regression · 599 cases needed at 80% power

The bet. The cheapest distribution available: a free tool used by exactly the people who buy the paid ones.

53JudgeEval pass rates corrected for judge bias

Your judge passes a third of the failures. Your pass rate is wrong.

70.4% → 62.4%
judge-reported pass rate, versus corrected for judge bias
Judge passes 43.3% of true failures · interval 48.7% to 73.1% from 300 human labels · v2 reads -0.9pp judged, -2.2pp corrected

The bet. Every team with an LLM judge reports a number that is wrong by an amount they have never measured, and the fix is one formula.

54SlicesEval gates that check every slice

The aggregate passed. Three customer slices did not.

0 → 3
regressions the aggregate gate saw, versus per slice after correction
Aggregate moved +0.5pp and passed · de-enterprise fell -7.9pp · 1 uncorrected false failure removed

The bet. Every team with an eval gate has been burned by the aggregate at least once, and the fix is a short function nobody has shipped as a gate.

55RecallRecall measured against what the pool misses

Your retrieval recall is 98% against a pool your retrievers built

100% → 81.1%
recall against the labelled pool, versus against what the pool is missing
985 labelled relevant documents across 200 queries · estimated relevant population 1,215 (1,162 to 1,268) · BM25 60.7% and dense 51.9% against it

The bet. Every RAG team has a recall number that never moves, and the engineer who owns it already suspects why.

56ConsensusFind the annotators whose labels are noise

Firing the bottom five annotators fires the wrong one

5 → 4
annotators flagged by raw agreement, versus real after shrinkage and correction
1,570 labels across 1,435 items to re-check · bar is 87.8% pooled agreement against 52.0% by chance · the rule's extra victim had 32 items

The bet. Every labelled dataset has a bottom-five rule and none of them have a test; the re-label order is a deliverable the vendor cannot write.

Services & capital

4 products
57Reliability SprintTwo weeks, fixed price, agent made shippable

Your agent works in the demo and nobody can say whether it is safe to ship

2.5 → 9.3
reliability score out of ten, week one versus handover
Weakest dimension on arrival: escalation calibration · blended half on the mean and half on the minimum, so a critical gap cannot be averaged away

The bet. Revenue in week two, a domain you actually know, and three customers writing the product spec instead of a guess.

58DiligenceAudits the numbers before you wire

You are about to wire on a retention chart nobody checked

5% → 51.2%
stated false-positive rate, versus the real one after 14 looks
The headline experiment was stopped on the fourteenth look · 41% of the retention cohort has not been around long enough to churn

The bet. Funds pay quickly, decide fast, and have nobody who checks the statistics behind the growth claims.

59Eval Build-OutTwo weeks to an eval suite you can trust

You have 62 test cases and you know that is not enough

62 → 840
eval cases, what they had versus what the effect size requires
Scorecard 4.0 to 8.0 out of ten · 4 duplicated prompts and 3 unreportable slices removed

The bet. The shortest path in this entire portfolio from a warm introduction to a paid invoice.

60UnderwritePrices an accuracy SLA you can actually back

Your customer wants an accuracy guarantee and nobody can price one

$0 → $102,375
annual premium priced off the average, versus off the tail
Mean accuracy 97.2% sits above the 95.0% guarantee, so an average-based quote sees no risk · 14 of 240 weeks breached anyway

The bet. Genuinely nobody is doing this, and the reason is that it is hard rather than that it is unwanted.

How to read this

The point is not to build 60 companies.

It is to stop betting a year on one guess. 60 landing pages, 60 pitches, one shared engine, and a signup form on each. Which one people give you their email for is a better signal than which one felt most exciting to build.

Everything here is a prototype. Real math on synthetic data with a known injected ground truth, real positioning, real capture. No auth, no billing, no customer data. The one that gets traction is the one that earns a backend.