Case studies
Banking

AML Transaction Monitoring When 95% Is Noise

Triage built to clear half to two-thirds automatically, before anyone opens a file.

CreateOS for AML Transaction Monitoring
On this page

At a Glance

MetricBeforeAfter
Alerts investigated per year500,000500,000 triaged, 191,000 to 263,000 reaching a human
False positives95%+ of all alerts50% to 65% removed before a person sees them
Fully loaded cost per alert$50$25 to $30
Annual alert investigation cost$25.0M$12.5M to $15.0M
Analyst hours spent clearing noiseBaseline79,000 to 103,000 hours returned
Investigation time, escalated case~2 hoursDown 50% to 60%
Time to detection, genuine activityBaselineDown 35% to 50%
Catch rateBaselineNo degradation, proven in shadow mode
Audit coveragePartial100% of decisions replayable, including auto-clears
Projected. Modeled on stated assumptions and published sources, not measured from a delivered deployment.

$10M to $12.5M per year in cost reduction on a $25M base.

The entire saving is concentrated in analyst capacity returned to genuine risk.

Challenge

Its rules and thresholds generated roughly 500,000 alerts a year, over 95% of them false positives, which is the industry norm at institutions of this size. Every one still landed in a queue for an analyst to open, research, and write up as a non-event.

  • $25 million a year, mostly spent documenting non-events. Roughly $50 in fully loaded cost on every alert touched, across 500,000 alerts.
  • 20 minutes for routine triage, two hours for a SAR case. And past twenty hours for the complex ones, with analysts working in permanent backlog hundreds of cases deep.
  • The industry response for a decade has been to hire. Labour is around 57% of financial crime compliance spend, and at a bank this size personnel run at 50% to 60% of the AML budget against 30% to 40% for technology. The queue still never empties.
  • Noise does not just cost money, it buries signal. The genuine suspicious activity was still in the pile, and every hour spent clearing a false hit is an hour not spent finding it.
  • The examiner asks about the clears, not just the catches. Whether every decision, including every decision that nothing was wrong, was reasoned and documented.
  • Previous pilots died in procurement, not the proof of concept. Information security would not approve a third-party platform processing transaction data outside the bank's boundary, and the second line would not accept a black-box score it could not explain to a regulator.

What the queue costs before anything is found

500,000

Alerts a year

~$50

Fully loaded cost per alert touched

$25.0M

A year, over 95% of it spent documenting non-events

Labour is roughly 57% of financial crime compliance spend, and the decade-long industry answer has been to hire against the queue. The cost is the visible half; the buried signal is the other.

Solution

  • Every alert gets triaged, and noise clears with a documented rationale. Genuine risk is enriched, classified, and escalated to a human as a near-complete case, so the analyst's day stops being a queue to empty.
  • Evidence assembly happens before a person opens the case. KYC profile, transaction history, counterparties, and adverse media consolidated in parallel rather than pulled from three systems by hand.
  • Activity is classified against known typologies. And against the bank's own scenario library, so the disposition carries a typology rationale rather than a score.
  • Separation of duties is a defensibility requirement. An examiner needs a clean split between gathering evidence, classifying risk, and reaching a conclusion. That is why it is many narrow agents rather than one model.
  • Analysis runs inside the bank's boundary. Control plane and storage sit in the bank's region, each investigation in its own guest kernel, with egress allowlisted in the kernel so transaction data cannot leave.
  • The filing decision stays with a human. Permanently and by design. The agents assemble, classify, and draft. The person decides.

The same 500,000 alerts, three ways

  • 95%

    Are false positives

    The industry norm at an institution of this size, and the reason the queue exists.

  • 50-65%

    Clear before a person sees them

    Each with a written rationale, not a score. This is the committed range.

  • 100%

    Replayable, including auto-clears

    The examiner asks about the clears, not only the catches.

The first ring is the problem and the second is the commitment against it, which leaves 191,000 to 263,000 alerts still reaching a person. The catch rate is held flat and proven in shadow mode before any of this binds.

What Made It Approvable

The reason this cleared information security and the second line of defense, when previous pilots had not, comes down to the layer underneath the agents, which most agent vendors rent and cannot control.

Every investigation runs in a Firecracker micro-VM with its own guest kernel, so analysis code executes against sensitive transaction data inside a contained boundary with a hard blast radius. Outbound network access is allowlisted in the kernel using eBPF, which means the enrichment agent can reach approved adverse-media providers, watchlists, and corporate registries and nothing else. Transaction data cannot be exfiltrated because the network path does not exist. The control plane is fully self-hosted, which is what makes the bank the owner of its model risk governance under the OCC's 2026 expectations rather than a tenant on someone else's black box. CreateOS is SOC 2 Type II and ISO 27001 certified.

The system also forks. One configured investigation environment forks per alert, so the bank's entire daily alert volume is processed in parallel rather than queued. Cases waiting on more information or on a human are paused rather than left burning compute.

The Guardrail That Mattered Most

In AML the failure mode that ends careers is the false negative, not the false positive. Aggressive noise reduction that risks missing genuine suspicious activity is the one outcome no risk officer will accept, and no honest vendor should promise.

So the system ran in shadow mode for four weeks before it cleared a single alert for real, triaging in parallel with the human team and making no binding decisions. Agent output was compared to human output case by case, and the gate for go-live was not the size of the false positive reduction. It was proof that the reduction came with no missed suspicious activity.

Auto-clear was then switched on for the lowest-risk, highest-noise alert types first, with full audit logging and human spot-checks, and the scope widened as the evidence built.

The gate for go-live was never the size of the reduction.

It was proof that the reduction came with no missed suspicious activity.

Outcome Derived

Of 500,000 alerts a year at $50 fully loaded, roughly 475,000 (about $23.75 million) are noise.

  • $10M to $12.5M removed a year. At the conservative committed 40% to 50%, with blended cost per alert falling from $50 to between $25 and $30. A sustained 60% takes it to roughly $20 and $15 million. We commit to the floor and treat the rest as upside.
  • 79,000 to 103,000 analyst hours returned. Clearing 50% to 65% of false positives removes 237,000 to 309,000 investigations from the queue. The bank did not cut the team, it redeployed onto genuine risk, which is what the regulator rewards and what the analysts wanted.
  • Escalated investigation time down 50% to 60%. The analyst opens a case already enriched, classified, network-mapped, and drafted rather than a bare alert, and SAR narrative drafting goes from hours to minutes.
  • Time to detection improved 35% to 50%. Which matters as much as the cost line, because catching real activity faster is what an examiner actually cares about.
  • 100% of decisions replayable, including every auto-clear. That is the deliverable, not a byproduct. A programme that cannot explain why it cleared something is not one an examiner will accept, however good its numbers look.
  • Regulatory exposure as context, not a promise. Fines totalled $3.8 billion in 2025 and single actions run into the hundreds of millions. We put no dollar figure on avoided fines, because that number is unprovable and a serious risk officer will discount an entire business case that tries.

What We Would Prove, and How

The engagement starts with the bank's own baseline, not ours, and proving no false negatives is treated as the entire game, not a formality.

  • Phase 0, weeks 1 to 2. Baseline: cost per alert, false positive rate, investigation time, and current catch rate, measured against the bank's own numbers. This baseline becomes the contract's yardstick.
  • Build and integration, weeks 2 to 6. Connect to the monitoring system, KYC store, transaction data, and enrichment providers along allowlisted paths.
  • Shadow mode, weeks 6 to 10. Longer than in onboarding, on purpose: in AML, proving the system did not introduce false negatives is the entire game.
  • Controlled go-live, from week 10. Auto-clear switches on for the lowest-risk, highest-noise alert types first, with full audit logging and human spot-checks, and the scope widens as the evidence builds.

Success criteria are agreed up front against the Phase 0 baseline: false positive reduction of at least 50%, no degradation in catch rate, cost per alert down at least 40%, and 100% audit coverage of every decision including auto-clears.

Highlights

  • Removes $10M to $12.5M a year in AML alert-investigation cost on a $25M base, with the saving concentrated entirely in analyst capacity returned to genuine risk.
  • Clears 50% to 65% of false positives before a human ever sees them, out of the 95%+ of the bank's 500,000 annual alerts that are noise.
  • Returns 79,000 to 103,000 analyst hours a year, redeployed onto genuine risk instead of clearing noise.
  • 100% of decisions are logged and replayable, including every auto-clear, built for the OCC's April 2026 model-risk expectations.
  • Self-hosted inside the bank's own boundary: control plane, storage, and Firecracker micro-VMs stay inside the bank's region, with eBPF-allowlisted egress so transaction data cannot leave.

Give Us One Stuck Pilot.

We'll have it in governed production before your next board meeting.