FDR control for the AML alert queue.
Most of the queue was never going to be a filing. Someone still has to read all of it.
Legacy transaction monitoring fires hundreds of rules across hundreds of customer segments every day with no correction for how many tests that is, so the queue fills with noise by construction. AML Triage ranks it by statistical evidence, flags segment behaviour no rule was written for, and logs every suppression for the examiner.
Fourteen real series, one injected typology, detection computed live in your browser.
The problem
The false positive rate is not a tuning problem. It is what the design produces.
A rules-based transaction monitoring system is a large set of independent thresholds, each one written for a typology somebody already thought of. Every rule is an independent test, evaluated fresh against every customer segment, every day. Nothing anywhere in that pipeline accounts for the number of tests being run, so the volume of alerts that close with no filing is a property of the arithmetic rather than a sign the thresholds were set badly. Tuning a rule moves which noise you get. It does not change how much.
The insight
You are already running tens of thousands of tests. Start correcting for it.
Two problems fall out of the same fact. Rules only encode typologies somebody wrote down, so anything novel passes through untouched. And because each rule is an independent test with no multiplicity control, the false positive volume scales with the size of the rule book rather than the size of the risk. Benjamini-Hochberg control across the whole alert population attacks the first problem by ranking on evidence instead of rule identity. Changepoint detection on segment-level behaviour attacks the second by looking for regime shifts directly, which does not require anyone to have anticipated the pattern. To be explicit about what this is not: AML Triage does not decide whether to file a SAR, does not replace analyst judgement, and does not touch the obligations the institution already has. It orders the queue and marks low-evidence alerts as low-evidence, with the q-value and the input series kept for every one.
CUSUM on day-over-day change for sustained shifts, Bayesian Online Changepoint Detection (Adams & MacKay 2007) for abrupt regime changes in segment behaviour, and Benjamini-Hochberg FDR control across every rule × segment combination in the run. Every alert, surviving or suppressed, keeps its q-value, its detector trace, and its input series in an immutable record.
How it works
Four steps, no data science team
Alerts and dispositions from the existing monitoring system, plus segment-level transaction aggregates. No rule changes, no replacement of the system of record, no new instrumentation.
Each alert is evaluated as one test among all tests run that day. Benjamini-Hochberg gives every alert a q-value that already knows how many hypotheses were in the batch.
Deposit velocity, cross-border share, average ticket, structuring indicators, per segment. A typology with no rule behind it still shows up as a regime change in the segment that is doing it.
Highest-evidence alerts first, with the series and the statistics attached. Suppressed alerts are still there, still dispositionable, still logged with the reason and the q-value that produced the ranking.
Who it is for
The BSA officer who has to explain the queue to an examiner
Compliance operations teams drowning in a queue they cannot staff their way out of. Usually 8 to 60 analysts, a legacy monitoring system nobody wants to rip out, and an examination cycle that makes any change to the alert process a documented decision.
Pricing
- –Full detection engine
- –Ranked queue with q-values
- –Segment changepoint detection
- –Exportable audit record
- –Unlimited segments and rules
- –Case management integration
- –Model documentation pack for validation
- –Suppression review workflow
- –SSO and immutable audit log
- –Self-hosted deployment
- –Model risk validation support
- –Custom typology detectors
- –Named compliance engineer
Competition
What exists, and what it does not do
| Who | What they do | The gap |
|---|---|---|
| NICE Actimize | The incumbent rules engine at most banks, with a machine-learning scoring add-on. | The scoring layer is still per-alert. Nothing in it accounts for how many rules fired across how many segments, so the multiplicity that generates the volume is untouched. |
| Unit21 | Modern no-code rules and case management, popular with fintechs and neobanks. | Makes it much easier to write and tune rules, which means more rules and more tests. Better ergonomics on the thing that causes the problem. |
| Hummingbird | Investigation and SAR filing workflow built around the analyst. | Starts after the alert exists. It makes reading the queue faster without changing what enters the queue. |
| Feedzai / Featurespace | Behavioural risk scoring, largely built for real-time payment fraud. | Optimised for per-transaction decisions in milliseconds. AML triage is a batch problem about the population of alerts, and the statistics that matter are the ones across the batch. |
The honest failure mode is procurement, not detection. Any layer that changes how alerts are prioritised is a model under SR 11-7, which means independent validation, documentation, and a conversation with an examiner before it goes anywhere near production. That is a nine to eighteen month cycle at a bank, and no BSA officer wants to be the first one explaining a statistical suppression layer to a regulator. If the model documentation pack is not good enough to survive validation on its own, the detection quality does not matter. There is also a real ceiling: this ranks and explains, it never decides to file or not file, so the analyst headcount it saves is bounded by how much of the queue is genuinely reviewable faster rather than skippable.
Market
Priced against analyst headcount, not software budget
A US community bank or mid-size fintech runs a compliance operations team of 8 to 60 analysts. Fully loaded, one analyst is roughly $95k. The Program tier at $144k a year has to beat one and a half analysts of throughput to pay for itself, which is a much easier argument than replacing the monitoring system. There are around 4,000 US banks and credit unions plus several hundred licensed fintechs and neobanks with a BSA obligation. Eight hundred of them on Program is $115M ARR.