Why This Case Study Exists
The other three cases in this series describe what an agent workforce does to a bank's AML operation. It clears the false positive flood. It compresses investigation from hours to minutes. It surfaces the networks a single-alert view cannot see.
None of them are possible without this one.
A bank cannot auto-clear 285,000 alerts a year unless every single one of those clearances carries a rationale it can defend to an examiner. Speed and cost are what the chief financial officer buys. Defensibility is what the chief risk officer permits. The order of operations is not negotiable, and any AML automation conversation that starts with savings and treats the audit trail as a compliance appendix is a conversation that ends in procurement.
Transaction Monitoring That Survives the Second Line
Defensibility is not a fourth benefit sitting alongside savings. It is the precondition for all of them.
At a Glance
| Metric | Before | After |
|---|---|---|
| Decisions with a replayable rationale | Partial | 100%, including every auto-clear |
| Model risk ownership | Vendor's black box | The bank's, documented and self-hosted |
| Transaction data location | Third-party platform boundary | Bank's own boundary and region |
| Outbound data paths | Trust-based, contractual | Allowlisted in the kernel, enforced |
| Blast radius of a compromised agent | Platform-wide | One disposable micro-VM |
| Basis for auto-clear | Score | Written rationale, inspectable per alert |
| Certification posture | N/A | SOC 2 Type II, ISO 27001 |
| OCC 2026 model risk alignment | N/A | Bank owns the framework; MRM support underway |
Challenge
The bank had tried to automate transaction monitoring twice. Neither attempt failed on accuracy. Both failed in procurement, when the second line of defence asked three questions.
- Why did it clear that alert?. The vendor could produce a score, not a reason. Auto-clearing hundreds of thousands of alerts a year on a number it cannot explain is not an automated AML programme, it is an industrialised indefensible one.
- The examiner asks about the clears too. Not only whether the bank caught the bad actor, but whether every decision, including every decision that nothing was wrong, was reasoned and documented.
- Where is our transaction data?. On the vendor's infrastructure, in the vendor's region, protected by contract. Information security would not sign it, because a contractual promise is not a control.
- Whose model is this?. The question that ended both attempts, and April 2026 sharpened it. OCC Bulletin 2026-13 replaced detailed model risk expectations with higher-level governance principles, rescinded the 2021 interagency statement on BSA/AML model risk, and left generative and agentic AI outside its scope pending a future request for information. Less prescription from the supervisor means the bank, not its vendor, has to be able to show its own governance.
- The vendor keeps the model, the bank keeps the consent order. A bank running monitoring on a vendor's opaque model is a tenant carrying regulatory liability for a system it cannot inspect, document, or fully explain.
- Enforcement did not soften, it relocated. Fines totalled $3.8 billion in 2025, down from $4.6 billion in 2024: North America down 58%, EMEA up 767%, APAC up 44% (Fenergo, Global AML Fines Research Report 2025). Financial crime compliance costs borne by financial institutions were put at $206.1 billion globally by LexisNexis Risk Solutions and Forrester Consulting in their True Cost of Financial Crime Compliance study (2023).
- Two unacceptable positions. Keep paying an army of analysts to clear a queue by hand, or adopt automation the bank's own risk function would not defend.
The question that ended two attempts
- Where does the data sit?Vendor infrastructure, vendor region, protected by contract.
- Whose model is this?OCC Bulletin 2026-13 sharpened the question in April.
- Who carries the consent order?The vendor keeps the model. The bank keeps the liability.
Why did it clear that alert?
A score is not a reason, and the examiner asks about the clears too.
- Enforcement relocated, it did not soften$3.8B in fines in 2025.
- Two unacceptable positionsAn army of analysts, or automation nobody will defend.
- Auto-clear needs a written reason285,000 a year, each defensible years later.
Solution
- Every decision is reasoned, including the clears. Not a score or a confidence interval, but the evidence gathered, the typology tested, the relationships examined and dismissed, and the written rationale for the disposition. An auto-clear is a documented conclusion, inspectable years later.
- The bank owns the model. Control plane and storage are fully self-hosted inside the bank's own infrastructure and region, so the model risk function can inspect, document, and govern the system as its own. Under a principles-based supervisory framework, that is the position no tenant on a vendor platform can occupy.
- We say where the tooling stands rather than trade on a roadmap. Self-hosting makes that ownership real today. Dedicated model risk management tooling aligned to the framework is under active development.
- The data cannot leave, and that is enforced not promised. Egress is allowlisted in the kernel, so enrichment reaches approved adverse-media providers, watchlists, and registries and nothing else. Not a policy saying data will not be sent somewhere, but an architecture in which it cannot be.
- The blast radius is one disposable machine. Every investigation runs in its own guest kernel and is destroyed when the case closes, with the VM lifecycle as the kill path, so a misbehaving agent is suspended rather than argued with.
- The human keeps the regulated decision. The decision to file stays with a human investigator, permanently and by design. That is a product constraint, and it is the answer a board risk committee always asks for.
- Most vendors own the logic and rent the runtime. Which is why they can promise a rationale but not a boundary, or a boundary but not an inspectable model. We own both layers.
Outcome Derived
Defensibility is not a fourth benefit sitting alongside the others. It is the precondition for all of them.
- The other three cases become approvable. The $10M to $12.5M of false-positive reduction, the $5.4M to $6.5M of investigation compression, and the network detection uplift are all unreachable for a bank whose risk function will not sign the system off.
- 100% audit coverage, including auto-clears. When an examiner selects an alert from three years ago and asks why it was cleared, the answer is a chain of reasoning rather than a number.
- Model risk ownership, not tenancy. The bank documents and governs the system as its own model, so it is not carrying regulatory liability for a black box it cannot inspect.
- Procurement stops being where the project dies. Where does the data live and why did it clear that alert now have structural answers rather than contractual ones. Information security reviews an architecture in which the exfiltration path does not exist.
- Regulatory exposure as context, not a promise. We put no dollar figure on avoided fines. That number is unprovable and a serious risk officer will discount an entire business case that tries to use it. We name it because it is why this reaches the board at all.
What defensibility is worth
Cannot answer the question
- Nothing shipsTwo attempts already stopped in procurement, not in proof of concept.
- The bank carries the liabilityRegulatory exposure on a model it cannot inspect.
Answered structurally
- $10M to $12.5MFalse-positive reduction, from the alert-triage study.
- $5.4M to $6.5MInvestigation cost-out, from the case-assembly study.
What We Would Prove, and How
Defensibility is not demonstrated in a slide. It is demonstrated by handing the bank's own second line of defense the system and letting them attack it.
- The examiner simulation. Before go-live, the bank's model risk and compliance functions select alerts at random from the shadow-mode period, including auto-clears, and demand the rationale. Not a summary. The full chain: what was gathered, what was tested, what was dismissed, and why. If a single decision cannot be replayed to their satisfaction, the gate does not open.
- The security review. Information security is given the architecture, not a description of it. Kernel-level egress allowlisting is a control that can be tested by trying to violate it, and we invite that.
- The model documentation. We produce the artifacts the bank's model risk function needs to own and document the system under OCC Bulletin 2026-13, and those artifacts are the bank's, not ours.
Success criteria, agreed up front: 100% of decisions including auto-clears replayable to the second line's satisfaction, egress controls verified by the bank's own security testing, and model documentation accepted by the bank's model risk function before a single alert is auto-cleared in production.
Methodology and Sources
The bank in this study is an illustrative composite. Its experience of automation pilots failing in procurement rather than in proof of concept reflects a documented pattern across the sector, and it is the pattern this case study is written to address.
The April 2026 model risk changes are OCC Bulletin 2026-13, Model Risk Management: Revised Guidance, issued 17 April 2026 with the Federal Reserve and the FDIC, at occ.gov. It replaced detailed expectations with higher-level principles, rescinded OCC Bulletin 2011-12 and the 2021 interagency statement on BSA/AML model risk, and excludes generative and agentic AI from scope. The argument that the bank rather than the vendor must own the framework is ours, drawn from that change, not a requirement stated in the bulletin. The 2025 fines figure of $3.8 billion, the prior-year comparison of $4.6 billion, and the regional enforcement shift come from Fenergo's Global AML Fines Research Report 2025, at resources.fenergo.com. The $206.1 billion of financial crime compliance cost borne by financial institutions comes from the LexisNexis Risk Solutions and Forrester Consulting True Cost of Financial Crime Compliance Global Study (2023), at risk.lexisnexis.com.
Platform claims in this document, specifically per-VM kernel isolation via Firecracker, kernel-level egress allowlisting via eBPF, full self-hosting of the control plane, production-hardened audit logging and role-based access control, and SOC 2 Type II and ISO 27001 certification, are statements about CreateOS as of July 2026. Model risk management support aligned to the OCC 2026 framework is in progress and is stated as such.
Sources: OCC (Bulletin 2026-13, Model Risk Management: Revised Guidance, 17 April 2026), Fenergo (Global AML Fines Research Report 2025), LexisNexis Risk Solutions and Forrester Consulting (True Cost of Financial Crime Compliance Global Study, 2023), FinCEN (SAR FAQs, 9 October 2025).
Highlights
- 100% of decisions carry a replayable rationale, including every auto-clear.
- Model risk ownership sits with the bank, documented and self-hosted, not a vendor black box.
- Transaction data stays inside the bank's boundary; outbound paths are allowlisted in the kernel.
- Blast radius of a compromised agent is one disposable micro-VM.
- Defensibility is the precondition that makes the other three AML cases approvable.



