Evidence for when a model quietly changed.
"The model got worse this week" is a feeling. Give it a date.
Providers change what sits behind a version string, whether that is quantisation, routing, or the serving stack, and your product changes with it. Drift Radar runs a frozen suite against every pinned model on a schedule and tells you, statistically, the day something moved.
Fourteen model-capability series across three vendors. The regression is in one of them.
The problem
You built on an API whose behaviour is not versioned the way you assumed.
A pinned model name is a promise about the weights, not about the whole serving path. Quantisation, speculative decoding, routing, and safety-layer changes all move observable behaviour without moving the version string. When it happens, every team on that model notices something at once, argues about whether it is real, and has no shared evidence. The complaint is loud, recurring, and completely unevidenced.
The insight
Freeze the tasks and the model becomes the only thing that can move.
Run a fixed suite at temperature zero, on a schedule, against pinned models, and the score is a time series whose only source of variation is the provider. That converts an argument into a changepoint problem, and a changepoint problem has an established answer. The reason this has not existed is not that the statistics are hard; it is that everyone runs their own evals privately, on tasks that change, without a schedule, which destroys the comparability the test needs.
A frozen task suite run three times daily at temperature zero against every pinned model. Per-capability scores, output length, and latency become series. CUSUM and Bayesian Online Changepoint Detection run per model × capability, with Benjamini-Hochberg FDR control across the full grid. Tracking six vendors across eight capabilities is 48 simultaneous tests, and without correction you would announce a false regression most weeks.
How it works
Four steps, no data science team
Every tracked model, every capability, every day, with the detected changepoints marked. Free because the dashboard is the distribution, and because a shared public record is worth more than a private one.
Aggregates hide exactly the regressions that hurt. A model that stays flat overall while losing eight points of schema adherence has broken every product that parses its output.
Tell it which models and which capabilities your product actually relies on, and get a message the day one of them moves, not the week your users start complaining.
The public suite is generic. Paid accounts run your tasks on the same schedule and the same statistics, so drift is measured against what you actually ship.
Who it is for
Teams whose product is downstream of somebody else’s model
Anyone whose product sits on a provider API and has been surprised by it. The free dashboard is for everyone; the paid product is for teams whose revenue depends on one capability of one model.
Pricing
- –Full public dashboard
- –All tracked models
- –Historical changepoints
- –Open methodology
- –Alerts on your models
- –Capability-level targeting
- –Slack and webhook
- –Early access to findings
- –Custom task suite
- –Runs on your prompts
- –Regression reports
- –Vendor escalation evidence
Competition
What exists, and what it does not do
| Who | What they do | The gap |
|---|---|---|
| Public leaderboards | Rank models against each other at a point in time. | Built to compare models, not to watch one model over time. A regression on a pinned model is not what they are looking for. |
| Your own eval suite | The right instinct, and what careful teams do. | Run irregularly, on tasks that change as the product changes, with no test applied to the result. That destroys the comparability that makes drift detectable. |
| Provider status pages | Report outages and incidents. | A behaviour change is not an incident. Nothing was down. |
| Arize / Langfuse | Observability over your own traffic. | Your traffic changes for your own reasons too. Isolating the provider requires holding the tasks fixed, which only a dedicated benchmark can do. |
Two real ones. First, providers may object to being publicly measured, and the relationship is asymmetric, though the suite uses ordinary paid API access and publishes its methodology, which makes the finding checkable rather than accusatory. Second, this may be a great distribution engine attached to a small paid product: everyone reads a free dashboard, and far fewer pay for alerts. That is why the private-suite tier exists, and it is the tier that has to be proven.
Market
Small paid market, unusually large audience
The honest sizing: the free dashboard could reach most engineers building on model APIs, and the paid conversion on that is thin. The revenue case rests on the private-suite tier sold to companies where one capability of one model is load-bearing for the product. A thousand of those at the Alerts tier plus two hundred private suites is $6M ARR. Modest, with distribution that makes everything else easier to sell.