At a Glance
| Metric | Before | After |
|---|---|---|
| Screening false-positive rate | 42% average, up to 95% at scale | Volume halved |
| Noise alerts reviewed by humans | 252,000 to 570,000 a year | Cut by half |
| Analyst hours on noise | 84,000 to 190,000 a year | 42,000 to 95,000 returned |
| Fraud and true-hit detection | Baseline | Up 45% to 61% |
| Manual share of KYC tasks | About two thirds (Fenergo, 2025) | Down 78% in processing time |
| Audit coverage of clear decisions | Partial | 100% replayable |
Direct labour recovered: $1.9M to $4.3M a year.
Analysts stopped reviewing noise. The alerts that reach a human are now alerts worth a human.
Challenge
Ask any KYC analyst what they actually do all day and you will not hear investigate financial crime. You will hear clear alerts that were never going to be anything.
- Roughly 600,000 alerts a year, around 42% false positives. Off sanctions, PEP, and adverse-media checks, each still opened, checked, documented, and closed. Twenty minutes, give or take. Then again.
- Screening teams at the largest institutions run into the thousands. Average annual spend on AML and KYC operations stands at $72.9M per firm (Fenergo, 2025 Financial Crime Industry Trends, a survey of 600 senior decision-makers). Automation of periodic KYC reviews averages roughly a third across respondents, so about two-thirds of the work is still done by hand.
- The four hundred and first alert does not get fresh eyes. This is the part that should worry a regulator more than it worries most banks. Noise does not just cost money, it erodes the attention that finds the real hit.
- Queues grow faster than headcount. Attrition in financial-crime teams is high, and the reason people give on the way out is the tedium.
- Rules-based tuning trades recall for precision. Tighten the matching logic and false positives drop, but so does detection, and no compliance officer will accept a change that might miss a true sanctions hit.
Why the tuning dial is the wrong control
Tighten the matching logic
- False positives fallFewer alerts reach the queue.
- So does detectionNo compliance officer signs off on a change that might miss a true hit.
Absorb the noise instead
- Volume halvedEach clear carries a written rationale, not a suppression rule.
- Detection up 45% to 61%Alerts are worked, not filtered out before anyone sees them.
Solution
- Clean files clear straight through, with the rationale recorded. Anything the system is not confident about routes to a named human, so analysts see exceptions rather than volume.
- Documents are read and verified rather than keyed. Identity documents, proofs of address, and income records extracted and authenticated, with the untrusted-file handling in the tightest boundary in the system.
- Screening noise is absorbed before it reaches a person. Sanctions, PEP, and adverse-media alerts are worked the way an analyst would work them, so what arrives at the queue is worth a human.
- Beneficial owners are resolved and screened in parallel. One configured environment forks per person, so a corporate structure unfolds all at once rather than one analyst at a time.
- Every decision carries a replayable rationale. Including the clears, logged as the work happens rather than reconstructed on request.
- Customer data stays inside the bank's boundary. Control plane and storage in the bank's own region, each case in its own guest kernel, with egress allowlisted in the kernel to approved list and registry providers and nothing else.
- The regulated decision stays with a human. Every genuine hit and every low-confidence case goes to a named reviewer by policy. CreateOS is SOC 2 Type II and ISO 27001 certified.
Outcome Derived
Analysts stopped reviewing noise. The alerts that reach a human are now alerts worth a human.
- $1.9M to $4.3M a year of direct labour recovered. Stated as a model so every input can be replaced with the client's own actuals.
- Detection improves rather than degrades. True-hit detection goes up 45% to 61%, because the agent absorbs noise rather than suppressing matches. That is the opposite trade from rules-based tuning.
- The engine inside the headline KYC number. False-positive triage is the single largest driver inside the 40% to 48% KYC cost reduction we commit to, which is why it is worth naming separately in a compliance conversation.
- The human stays in the sanctions decision. Every genuine hit, and every alert the agent is not confident about, still goes to a named reviewer by policy.
What We Would Prove, and How
- Weeks 1 to 2, baseline. Measure the bank's actual alert volume, false-positive rate, average review time, and true-hit rate. This is the yardstick the contract is written against, and it protects both sides in procurement.
- Weeks 2 to 6, build and integrate. Stand up the agent crew, integrate to the bank's screening providers and customer data along allowlisted egress paths, deploy self-hosted inside the bank's boundary.
- Weeks 6 to 8, shadow run. The agent processes live alerts in parallel with the human team without making binding decisions. Every agent clear is compared against the human clear. Recall is the metric that gets watched, not just precision.
- Week 8 onward, controlled go-live. Pre-clearing switched on for the alert categories where the shadow run showed no true-positive loss, expanding as the audit record builds.
Success criteria, agreed up front: false-positive volume reaching a human halved, zero true-positive loss against the shadow-run baseline, 100% audit coverage of every automated clear.
Highlights
- Halves screening false-positive volume reaching humans, recovering $1.9M to $4.3M a year in direct labour.
- Returns 42,000 to 95,000 analyst hours from noise review (21 to 47 trained analysts).
- True-hit detection improves 45% to 61% because the agent absorbs noise rather than suppressing matches.
- Recall stays where the compliance team set it; the agent does not tighten matching logic.
- 100% of automated clears are replayable; every genuine hit still goes to a named reviewer.
Frequently asked questions
How is sanctions screening false-positive clearance audited?
Every automated clear carries a replayable rationale, and 100% audit coverage of automated clears is a success criterion agreed before go-live. Genuine hits, and any alert the agent is not confident about, still go to a named reviewer by policy. The sanctions decision itself stays with a person.
Does cutting screening false positives mean tightening the matching rules?
No, and that is the point. Rules-based tuning trades recall for precision: tighten the logic and false positives fall, but so does detection. Here the agent absorbs the noise instead of suppressing matches, so recall stays where the compliance team set it. Zero true-positive loss against the shadow-run baseline is a stated success criterion.
How much analyst time does false-positive triage give back?
This blueprint models 42,000 to 95,000 analyst hours returned a year, the equivalent of 21 to 47 trained analysts, on a book of roughly 600,000 alerts at around 42% false positives. Those are modelled figures. The bank's actual alert volume, false-positive rate, review time, and true-hit rate are measured in Phase 0.
Where does alert data go when the agent enriches a screening hit?
Nowhere outside the bank. Control plane and storage run inside the bank's own boundary, each case runs in its own guest kernel, and egress is allowlisted in the kernel to approved screening providers and nothing else. Data cannot be sent elsewhere, because the network path does not exist.



