Pricing consistency: Reproducible by construction
Rating-factor coverage: 100% of decisions logged with factors
Decision replay: Replayable from the audit log
Bias testing: Testable factor by factor
Challenge
Insurance pricing is rate-regulated, and algorithmic underwriting is drawing the kind of scrutiny fair lending has drawn for years. An examiner does not ask whether the decision was quick. They ask why this risk was priced this way, whether a comparable risk was priced the same, and whether any factor is acting as a proxy for something the carrier may not price on.
- A hundred underwriters pricing by hand produced inconsistency. Two similar risks landing on two different desks came back with two different prices, and the carrier could not fully reconstruct why.
- That inconsistency was expensive in two directions. It showed up in the loss ratio as mispriced risk, on a $500M book where a single point is $5M, and as exactly the unevenness a fair-pricing regulator probes.
- Inconsistency is indistinguishable from arbitrariness. When you are the one being examined.
- The decision trail was partial. Rating factors lived in the underwriter's head or a note in the file, so reconstructing a two-year-old decision meant finding the underwriter and hoping they remembered.
- Every vendor solution made this worse. They wanted the filed rates, the rating model, and policyholder data in their cloud, and returned a price. The carrier answers the examiner, not the vendor, so a black-box price on someone else's infrastructure is a liability it could not discharge.
- Fast, cheap, and biased is the worst outcome, not the best. It is a market-conduct action, and a bigger number than any efficiency saving on the page.
What rate regulation is actually asking
Fast and cheap
- A market-conduct actionFines, remediation, rate refiling. On no ROI model anywhere.
- Inconsistency reads as arbitrarinessTwo similar risks, two desks, two prices, and no trail.
Reproducible and explained
- Every decision carries its factorsLogged as it happens, replayable two years later.
- Forked from one template per riskThe same logic on every submission, testable factor by factor.
Solution
CreateOS built the defensibility into the infrastructure rather than bolting it on as a report.
Could she defend the price?
A hundred underwriters pricing by hand produced two prices for two similar risks.
100%factors logged
A document, not a memory
Every decision replayable with its rating factors, years later.
No loss-ratio gain is claimed before champion-challenger testing on the carrier's own book.
- Reproducibility by construction. The environment is baked as a template and forked per submission, so every risk runs through identical logic. There is no different desk to land on, which answers both the inconsistency an examiner probes and the mispricing the loss ratio punishes.
- Separation of concerns is an audit control. A monolithic model that ingests a submission and emits a price is a black box that cannot be tested factor by factor. Narrow agents with clear responsibilities produce a chain that can be examined at each link.
- Every pricing decision carries its rating factors. Logged step by step as it happens, with every appetite check applied. When the examiner asks why this risk was priced this way, the answer is a document, not a search for the underwriter who made the call.
- Model ownership stays with the carrier. The rating model never leaves the boundary, so it stays the carrier's to own, document, and defend, in the hands of the party actually accountable for it.
- Pricing models cannot be exfiltrated. Egress is allowlisted in the kernel to approved data providers, so even a compromised or misbehaving agent cannot move them, because the path does not exist.
- Per-VM isolation on untrusted attachments. Every submission runs in its own guest kernel, and a misbehaving agent is suspended instantly through the VM lifecycle.
- Pricing authority and the bind stay with the underwriter. Out-of-appetite, referred, and complex risks route to a person. The agent assesses and recommends. That is a policy control, not a fallback.
CreateOS is SOC 2 Type II and ISO 27001 certified, so the procurement and information-security gates are cleared before the actuarial conversation starts.
Outcome Derived
The carrier got a pricing decision it could defend, and the consistency paid for itself.
| Metric | Before | After |
|---|---|---|
| Pricing consistency | Varies by underwriter | Reproducible by construction |
| Rating-factor coverage | Partial, often in the file note | 100% of decisions logged with factors |
| Decision replay | Find the underwriter, hope they recall | Replayable from the audit log |
| Bias testing | Not systematically possible | Testable factor by factor |
| Model ownership | At risk in vendor clouds | Stays inside the carrier's boundary |
| Data egress | Broad vendor access | Kernel-allowlisted, approved paths only |
| Certification posture | Vendor-dependent | SOC 2 Type II, ISO 27001 |
| Loss ratio | Baseline | Not claimed; tested champion-challenger first |
- The loss-ratio prize is the largest and most oversold number. A single point on the carrier's $500M book is $5M and three points is $15M, which dwarfs the $6M to $8M of operational saving. Which is exactly why we do not headline or guarantee it.
- It is proven the only way it honestly can be. Champion-challenger testing against the carrier's own loss experience, before a single policy is bound on agent pricing. A vendor putting a loss-ratio number on a slide before touching your book is telling you what they are willing to say, not what they are able to do.
- The downside appears on no ROI model. A market-conduct action for unfair or proxy discrimination is fines, remediation, rate-filing consequences, and a headline. Fair-pricing testing is a hard gate in the pilot because the entire value of the system is negative if the pricing is fast, cheap, and indefensible.
- Where the gap still is, stated plainly. Bias-testing tooling and formal model documentation support are still being built out. The certifications are done and self-hosting keeps the model in the carrier's hands, but an actuarial and market-conduct review will ask for systematic fairness evidence, and we would rather name that as an open engineering priority than let a client discover it in Phase 2.
What We Would Prove, and How
Weeks 1 to 2, baseline. Measure pricing consistency across underwriters on comparable risks, the current rating-factor coverage in the file, the loss ratio on the target line, and the carrier's existing fairness testing posture.
Weeks 2 to 6, build and integrate. Stand up the agent crew, integrate to policy administration, the rating engine, and approved data providers along allowlisted paths, encode appetite and filed rates, deploy self-hosted inside the carrier's boundary.
Weeks 6 to 10, champion-challenger shadow run. Agents price real submissions in parallel with the underwriters, binding nothing. Compare decisions and prices case by case, validate the loss-ratio claim against the carrier's own loss experience, and test explicitly for unfair and proxy discrimination. Proving the pricing is both better and fair is the whole gate. Nothing goes live until it clears.
Week 10 onward, controlled go-live. Straight-through quoting on the cleanest standard, in-appetite risks first, underwriters on every referral and decline, scope expanding as the champion-challenger evidence and the fairness testing build the record.
Success criteria, agreed up front: 100% explainable rating-factor coverage, reproducible pricing on comparable risks, no unfair or proxy discrimination, no degradation in loss ratio, and a validated loss-ratio improvement measured on the carrier's own book rather than promised from a benchmark.
Highlights
- Pricing is reproducible by construction. Every risk runs through an identical copy of identical logic, so two comparable risks are assessed the same way.
- 100% of decisions logged with their rating factors, replayable from the audit log rather than reconstructed from an underwriter's memory.
- The loss-ratio prize is real and we do not guarantee it. A single point on a $500M book is $5M, but it is a model-quality outcome proven only by champion-challenger testing against the carrier's own loss experience.
- A market-conduct action is not an efficiency miss. It is fines, remediation, rate-filing consequences, and a headline, which is why fair-pricing testing is a hard gate in the pilot rather than a later workstream.
- Bias-testing tooling and formal model documentation are still being built out. The certifications are done and self-hosting keeps the model in the carrier's hands, but systematic fairness evidence is an open engineering priority and we would rather name it here.



