Case studies
Banking

Finding the Money in Items Nobody Ever Opened

About 26,000 a year age out unexamined; that is where $4.8M of leakage sits.

CreateOS for Revenue Leakage Recovery
On this page

At a Glance

  • Client. Mid-size commercial bank, ~1.8M transactions per month, ~$1.2B net revenue
  • Starting point. 648,000 exceptions a year. Roughly 26,000 of them never properly worked before they aged out.
  • What CreateOS built. A six-agent reconciliation crew that works every exception, including the ones the clock used to eat
  • Leakage identified on the bank's own books. $4.8M annually
  • Recovered within claw-back windows. $3.1M in year one
  • Prevented from forming again. ~$1.7M a year
  • What the category's rule of thumb would have let us claim. $12M to $24M. We did not claim it.

Challenge

Every bank with a reconciliation function has a number it does not know. Not one it disputes or argues about in committee, but one it has genuinely never seen, because the process that would surface it is the process that runs out of time. This client's version was $4.8 million a year.

  • The team was exactly sized, not comfortably sized. 126 analysts on exception duty against 648,000 investigations a year at 22 minutes each. During the close, the queue consumed the entire team's capacity with no slack at all.
  • Roughly 26,000 items a year aged out unexamined. About 4% of annual exception volume. Some sat in suspense accounts, some were batch-closed with a generic disposition code because the close had to sign, some were written off under a materiality threshold: individually too small to fight about, collectively enormous.
  • The universal assumption is that they were timing differences. And most of them were. That assumption is not wrong, it is just not entirely right, and the part where it is wrong is where the money is.
  • Inside them sat real money nobody had looked for. Duplicate payments never clawed back. Fees never captured because the discrepancy that would have revealed the underbilling was never investigated. Interchange and chargeback recoveries that expired in their windows while the item sat in a queue.
  • The rule of thumb would have implied $12M to $24M. The figure this category quotes for leakage from manual reconciliation is 1% to 2% of net revenue. No named study stands behind it, and the variants that switch the denominator to payment volume are further adrift again. On $1.2 billion of net revenue, 1% to 2% is that range.
  • We refused to use it. Which is the whole point of this case study.

Two ways to size the same leakage

The category's rule of thumb

  • $12M to $24M1% to 2% of $1.2B net revenue, multiplied on a slide.
  • Nobody can check itNot measured on this bank's books, or on anyone's.

Opened and counted

  • $4.8M, measured0.4% of net revenue: less than half the bottom of the range.
  • $3.1M recovered in year oneEvidenced to an external auditor rather than asserted.
The gap is the useful part. Another bank might come in above 1% or at 0.2%, and nobody knows which until the aged items are opened. Every vendor in this category could have put $18M on a slide and asked for a share of it.

Solution

Every vendor in this category could have walked into that bank, multiplied $1.2 billion by 1.5%, put $18 million on a slide, and asked for a percentage of it. The percentage circulates widely enough that nobody in the room would have questioned it. It would have made a spectacular pitch.

It would also have been a lie, and the CFO would have known within about ninety seconds, because a bank's actual leakage is a function of its rail mix, counterparty quality, fee structure, materiality thresholds, and claw-back windows. It is not a function of the industry average. Quoting a portable percentage against a specific balance sheet is how you get walked out of a room you were winning. So we did the harder thing and went to find the real number.

  • Every aged item gets opened, not sampled. The capacity to work all 26,000 unexamined items is what turns an assumption into a measurement.
  • Each break is traced across systems to root cause. With the deterministic match kept rules-based and exact, because a guessed match is a misstated book.
  • Recoveries post through sanctioned paths only. Inside the bank's control thresholds, with anything above them routed to a named human.
  • Every finding is logged with its rationale. So a recovery can be evidenced to an external auditor rather than asserted.

The property that makes leakage recovery possible, rather than merely faster reconciliation, is fork-per-exception. One configured investigation environment forks for every single break. The queue does not exist any more. There is no line for an item to fall out of, no capacity ceiling for the close to slam against, and therefore no mechanism by which an exception ages out unexamined.

That is the actual unlock. Not that the agents investigate better than the analysts, though they arrive with the history already assembled rather than starting cold. It is that they investigate all of them. The 26,000 items a year that used to die in suspense now get worked, because working them costs a fork rather than an analyst-hour.

Then we ran nine weeks of shadow mode. The agents processed live exceptions in parallel with the human team, binding nothing, posting nothing. And in that window we did something the bank had never been able to do: we opened the items that historically would have aged out, and we found out what was actually in them.

Outcome Derived

Leakage identified on the bank's own books: $4.8 million a year, measured rather than modelled.

  • 22% of aged items carried genuine financial impact. Some 5,700 of roughly 26,000, at an average value at risk of around $840: duplicate payments never recovered, uncaptured and underbilled fees, chargeback and interchange recoveries expired in their windows.
  • The other 78% were exactly what the bank assumed. Timing differences and noise that resolved to nothing. The instinct was right about most of the pile and wrong about the part that mattered, with no way to know which without opening every item.
  • $3.1 million recovered in year one. Roughly 65% of the identified leakage was still inside a claw-back, correction, or recovery window and could actually be pursued.
  • ~$1.7 million a year prevented from forming again. The remaining 35% was real, identified, and time-barred. That is not a recovery, it is a diagnosis: the rate at which the bank was losing money to items it never got to, and it stops accruing.
  • 0.4% of net revenue against a 1% to 2% rule of thumb. Less than half the bottom of the range. We could have let $4.8 million stand alone as an impressive number. Instead: what we found on a real balance sheet was two and a half times smaller than what the category's rule of thumb would have entitled us to claim.
  • That gap is the most useful thing here. Your bank might come in above 1%, or at 0.2%. We do not know, you do not know, and anyone who claims to know before opening your aged items is selling you a number rather than finding you one.

How we priced it.

We did not price against the $4.8 million, and we did not price against the rule of thumb. The commercial structure is a value share on funds actually recovered, measured and agreed after the fact, sitting alongside a platform fee for the operational work. The bank pays a share of money that has demonstrably come back through the door.

This sits on top of, and entirely separate from, the $6.0 million to $8.0 million a year of hard operational cost reduction the same agent crew delivers by resolving exceptions 60% faster. We keep those two columns apart in every conversation. Cost reduction is a commitment. Leakage recovery is a validated outcome. A finance reviewer who finds a soft input smuggled into a hard column stops believing the hard column too, and they are right to.

Everything above was produced under the same controls as the rest of the system. The resolution agent can only post through sanctioned ledger and payment paths, allowlisted in the kernel using eBPF, so an agent chasing a recovery cannot touch anything outside policy because the network path does not exist. Recoveries above the bank's thresholds route to a named human. Every investigation, finding, and posting is logged with its rationale and is replayable for an external auditor. Zero unauthorized postings, enforced below the model rather than promised by it.

The Uncomfortable Question This Raises

If the bank had 26,000 items a year it was never opening, and 22% of them had money in them, then the bank had been losing roughly $4.8 million a year for as long as its exception queue had been capacity-bound.

Which is to say: for years. The leakage does not start when it is measured. Measuring it is simply the first time anyone has looked.

This is the part of the reconciliation conversation that nobody enjoys, and it is why leakage is a genuinely difficult thing to sell against. The CFO is not being asked to approve an improvement. They are being asked to find out something about their own books that they may prefer not to know, and that someone will eventually ask why they did not know sooner.

The honest framing, and the one that worked here: the number exists whether or not you look at it. It has been accruing this entire time. The only question on the table is whether it keeps accruing.

What We Would Prove, and How

Phase 0, weeks 1 to 2. Baseline. We measure exception volume, aged-item volume, current disposition of items that never get worked, existing suspense balances, and known write-offs. This establishes what the bank is currently not seeing.

Phase 1, weeks 2 to 6. Build and integrate. Stand up the agent crew, integrate to the general ledger, payment rails, and reconciliation tooling along allowlisted paths, encode matching rules and control thresholds, deploy self-hosted inside the boundary.

Phase 2, weeks 6 to 9. Shadow run and leakage discovery. The agents work every exception, including the ones that would historically have aged out, posting nothing. This is where the real leakage figure is found. It is measured against the bank's actual ledger, categorized, and split into recoverable and time-barred.

Phase 3, weeks 9 onward. Controlled go-live and recovery. Automated posting enabled for lowest-risk within-threshold items, recoveries pursued within their windows, everything above threshold routed to a human, full audit logging on.

Success criteria, agreed up front: every exception worked, zero items aged out unexamined, leakage identified and categorized against the bank's own ledger, recoverable portion pursued, zero unauthorized postings, 100% audit coverage.

Highlights

  • The other 78% were exactly what the bank had always assumed they were. Timing differences and noise that resolved to nothing. The bank's instinct was right about most of the pile. It was wrong about the part that mattered, and it had no way of knowing which part that was without opening every item.
  • The remaining 35% was real, and identified, and time-barred. Too old to recover. The value there is not a recovery, it is a diagnosis: this is the rate at which the bank was losing money to items it never got to, and from the moment every exception gets worked in time, that leak closes. It stops accruing.
  • The percentage this category quotes is 1% to 2% of net revenue. We came in at less than half of the bottom of that range.
  • We are putting that in the case study on purpose. We could have quietly not mentioned the rule of thumb and let $4.8 million stand on its own as an impressive number, which it is. Instead we are telling you that the number we found on a real balance sheet was two and a half times smaller than the number the category's own rule of thumb would have entitled us to put on a slide.
  • That gap is the most useful thing in this document. It is what a real measurement looks like against a circulating rule of thumb, and it is why we will not quote you one. Your bank might come in above 1%. It might come in at 0.2%. We do not know, you do not know, and anyone who tells you they know before they have opened your aged items is selling you a number rather than finding you one.

Give Us One Stuck Pilot.

We'll have it in governed production before your next board meeting.