Live demo

8 queues in. 1 out.

Eight queues, 40 weeks of ticket counts, a bot on billing from week 26, and a season that cut volume on every queue in the same week. Each queue is fitted against the others. The one where something happened has to be the one that got the bot.

Vendor dashboard
35.0%
conversations "resolved by AI"
Before/after
25.0%
billing volume, post against pre
Against the control
17.4%
159 tickets a week removed
Truth injected
18.0%
what the generator removed
passControl tracked billing before launch1.6%

Pre-launch RMSE as a share of weekly volume, across 26 weeks, against a tolerance of 4.0%. Pre-period correlation 0.959. Above the tolerance, nothing below is reported.

passSomething happenedp = 0.0020

Difference-in-differences against the scaled donor mean, permutation test over pre and post labels. The gap is sized only after this clears. DiD puts it at 17.7%, in agreement with the synthetic control.

passPlacebo on account-access, which got no bot1.9%

The same fit on a queue where nothing was launched, p = 0.598. It has to find nothing, and it does. Post gap against pre noise is 1.0x there and 10.2x on billing.

Every queue fitted against the others. Two are refused a number because no blend of the rest tracks them.
QueueRoleTracking errorControlDeflectionDiD p
billingtreated1.6%tracks17.4%0.002
shippingplacebo11.5%no controlnot reportedn/a
technicalplacebo2.1%tracks-1.5%0.673
account-accessplacebo2.4%tracks1.9%0.598
refundsplacebo3.4%tracks-2.0%0.666
returnsplacebo1.6%tracks3.2%0.236
onboardingplacebo1.9%tracks-2.8%0.242
integrationsplacebo68.6%no controlnot reportedn/a
Deflection before/after books
25.0%
Deflection measured
17.4%

Billing fell 25.0% after launch. Its control, a blend of account-access 48.2%, refunds 41.0%, returns 5.7%, technical 5.1%, fell 9.2% over the same 14 weeks with no bot on it. That is the season, and before/after booked it as deflection.

The vendor number is not a lie, it is a different quantity: conversations the bot answered after which the customer went quiet. Many of those would have gone quiet anyway. The number the renewal turns on is tickets that would have existed and do not, and that can only be measured against queues that got no bot. Where no such control exists, as for shipping and integrations here, this reports no number rather than a bad one.