Evidence for when a model quietly changed.

"The model got worse this week" is a feeling. Give it a date.

Providers change what sits behind a version string, whether that is quantisation, routing, or the serving stack, and your product changes with it. Drift Radar runs a frozen suite against every pinned model on a schedule and tells you, statistically, the day something moved.

No spam. One email when it is ready to try.

DetectedJSON schema adherence · vendor-a / pinned-2026-05changepoint 2026-08-12 · version string unchanged
88.2%93.4%98.6%changepointJun 29Aug 27

Fourteen model-capability series across three vendors. The regression is in one of them.

The problem

You built on an API whose behaviour is not versioned the way you assumed.

A pinned model name is a promise about the weights, not about the whole serving path. Quantisation, speculative decoding, routing, and safety-layer changes all move observable behaviour without moving the version string. When it happens, every team on that model notices something at once, argues about whether it is real, and has no shared evidence. The complaint is loud, recurring, and completely unevidenced.

One number
What most leaderboards report per model. A regression confined to one capability disappears into an aggregate.
No changelog
Serving-side changes behind an unchanged version string are frequently shipped without a published entry.
Every team, alone
Each customer re-derives the same suspicion in private with no way to confirm it against anyone else.

The insight

Freeze the tasks and the model becomes the only thing that can move.

Run a fixed suite at temperature zero, on a schedule, against pinned models, and the score is a time series whose only source of variation is the provider. That converts an argument into a changepoint problem, and a changepoint problem has an established answer. The reason this has not existed is not that the statistics are hard; it is that everyone runs their own evals privately, on tasks that change, without a schedule, which destroys the comparability the test needs.

Method

A frozen task suite run three times daily at temperature zero against every pinned model. Per-capability scores, output length, and latency become series. CUSUM and Bayesian Online Changepoint Detection run per model × capability, with Benjamini-Hochberg FDR control across the full grid. Tracking six vendors across eight capabilities is 48 simultaneous tests, and without correction you would announce a false regression most weeks.

How it works

Four steps, no data science team

01
A public dashboard, free forever

Every tracked model, every capability, every day, with the detected changepoints marked. Free because the dashboard is the distribution, and because a shared public record is worth more than a private one.

02
Detection per capability, not per model

Aggregates hide exactly the regressions that hurt. A model that stays flat overall while losing eight points of schema adherence has broken every product that parses its output.

03
Alerts on the models you depend on

Tell it which models and which capabilities your product actually relies on, and get a message the day one of them moves, not the week your users start complaining.

04
Bring your own suite

The public suite is generic. Paid accounts run your tasks on the same schedule and the same statistics, so drift is measured against what you actually ship.

Who it is for

Teams whose product is downstream of somebody else’s model

Anyone whose product sits on a provider API and has been surprised by it. The free dashboard is for everyone; the paid product is for teams whose revenue depends on one capability of one model.

Pricing

Public
$0
Everyone, forever
  • Full public dashboard
  • All tracked models
  • Historical changepoints
  • Open methodology
Most common
Alerts
$200/mo
Teams with a critical dependency
  • Alerts on your models
  • Capability-level targeting
  • Slack and webhook
  • Early access to findings
Private suite
$1,500/mo
Your tasks, your schedule
  • Custom task suite
  • Runs on your prompts
  • Regression reports
  • Vendor escalation evidence

Competition

What exists, and what it does not do

WhoWhat they doThe gap
Public leaderboardsRank models against each other at a point in time.Built to compare models, not to watch one model over time. A regression on a pinned model is not what they are looking for.
Your own eval suiteThe right instinct, and what careful teams do.Run irregularly, on tasks that change as the product changes, with no test applied to the result. That destroys the comparability that makes drift detectable.
Provider status pagesReport outages and incidents.A behaviour change is not an incident. Nothing was down.
Arize / LangfuseObservability over your own traffic.Your traffic changes for your own reasons too. Isolating the provider requires holding the tasks fixed, which only a dedicated benchmark can do.
How this fails

Two real ones. First, providers may object to being publicly measured, and the relationship is asymmetric, though the suite uses ordinary paid API access and publishes its methodology, which makes the finding checkable rather than accusatory. Second, this may be a great distribution engine attached to a small paid product: everyone reads a free dashboard, and far fewer pay for alerts. That is why the private-suite tier exists, and it is the tier that has to be proven.

Market

Small paid market, unusually large audience

The honest sizing: the free dashboard could reach most engineers building on model APIs, and the paid conversion on that is thin. The revenue case rests on the private-suite tier sold to companies where one capability of one model is load-bearing for the product. A thousand of those at the Alerts tier plus two hundred private suites is $6M ARR. Modest, with distribution that makes everything else easier to sell.

Get alerted on your models

No spam. One email when it is ready to try.

Or just go look at the demo first →