Live demo
8 queues in. 1 out.
Eight queues, 40 weeks of ticket counts, a bot on billing from week 26, and a season that cut volume on every queue in the same week. Each queue is fitted against the others. The one where something happened has to be the one that got the bot.
Pre-launch RMSE as a share of weekly volume, across 26 weeks, against a tolerance of 4.0%. Pre-period correlation 0.959. Above the tolerance, nothing below is reported.
Difference-in-differences against the scaled donor mean, permutation test over pre and post labels. The gap is sized only after this clears. DiD puts it at 17.7%, in agreement with the synthetic control.
The same fit on a queue where nothing was launched, p = 0.598. It has to find nothing, and it does. Post gap against pre noise is 1.0x there and 10.2x on billing.
| Queue | Role | Tracking error | Control | Deflection | DiD p |
|---|---|---|---|---|---|
| billing | treated | 1.6% | tracks | 17.4% | 0.002 |
| shipping | placebo | 11.5% | no control | not reported | n/a |
| technical | placebo | 2.1% | tracks | -1.5% | 0.673 |
| account-access | placebo | 2.4% | tracks | 1.9% | 0.598 |
| refunds | placebo | 3.4% | tracks | -2.0% | 0.666 |
| returns | placebo | 1.6% | tracks | 3.2% | 0.236 |
| onboarding | placebo | 1.9% | tracks | -2.8% | 0.242 |
| integrations | placebo | 68.6% | no control | not reported | n/a |
Billing fell 25.0% after launch. Its control, a blend of account-access 48.2%, refunds 41.0%, returns 5.7%, technical 5.1%, fell 9.2% over the same 14 weeks with no bot on it. That is the season, and before/after booked it as deflection.
The vendor number is not a lie, it is a different quantity: conversations the bot answered after which the customer went quiet. Many of those would have gone quiet anyway. The number the renewal turns on is tickets that would have existed and do not, and that can only be measured against queues that got no bot. Where no such control exists, as for shipping and integrations here, this reports no number rather than a bad one.