Why generic LLMs cannot manage commercial P&C risks

[4]

Last updated on

The real problem isn’t the model. It’s the idea that a model is enough.

Underwriting is being rebuilt around AI. The real question isn’t whether AI can help. Demos prove it every day. It’s what separates a demo from an intelligence capable of running an entire portfolio, continuously, and standing up to a regulator.

More than half of generative AI projects are abandoned after the proof-of-concept phase (Gartner). Most are building the wrong thing, because they confuse two logics. Consumer AI must be intuitive, demonstrative, conversational. Enterprise AI is expected to be deterministic, auditable, whilst maintaining marginal costs in check.

The press confirms this, more bluntly than an analyst firm would allow: “This gap between performance expectations and production reality is precisely why AI agents aren’t the one-size-fits-all tool the tech industry desperately wants them to be. Today’s reality is falling way behind the hype.” (Futurism, May 2026)

This observation is not isolated. It runs through the discussions of practitioners operating agents in production. The vocabulary of deterministic harnesses, of structural gates rather than prompt-based guardrails, of cost per outcome rather than per token, keep emerging with striking convergence.

While demos continue to impress, practical know-how is slowly taking shape in production. These concepts are not a Continuity invention. Rather they are establishing themselves as an industry practice. What Continuity brings to the table however is the underwriting precedent accumulated over seven years of portfolio analysis, which no generic architecture carries.

Meanwhile, thinking standard Consumer AI can deliver what a specialised Entreprise AI could, is the most costly mistake in AI for insurance in 2026.

Why generic LLMs fail the production test: three structural barriers

Barrier 1 | The cost of scale, and the false promise of cheap tokens

The real ambition is continuous intelligence across the entire portfolio: every policy, at every cycle, not a sample. The value only becomes real at scale.

The metric that matters is not cost per token. It is cost per outcome, at scale. And there is no single lever to control it: better or cheaper models, self-hosting, fine-tuning, better prompting, and many other approaches besides. Mastering cost per outcome is a continuous architectural endeavour, redefined with every generation of models, not a setting fixed once and for all.

The claim that “token costs have fallen by a factor of 300” is everywhere. It has become misleading. Not because it is false, but because it looks at the wrong metric. Gartner itself forecasts a drop of around 90% in inference cost by 2030, with intelligence trending towards zero cost. The problem is not the price of the token. It is that a generic LLM is priced per prompt, with no notion of outcome. Applied to a portfolio of 100,000 contracts, the cost structure becomes unsustainable. Not because the token is expensive, but because there is no concept of cost per outcome.

The frontier, meanwhile, remains volatile. Every quarter, the table is completely reshuffled: Kimi-K3 this month, GLM-5.2 before it, the Sonnet / GPT / Gemini reshuffles before that again. Providers get rid of their older models and force migration towards frontiers at higher unit cost; reasoning models consume 5 to 8 times more tokens per task. The lesson is not in the score of any given model in any given quarter. It is structural: the frontier is a moving target, and banking on a specific model is like betting on a market shifting beneath our feet.

This is precisely the warning from Gartner itself, in the forecast announcing commoditisation: “CPOs who mask architectural inefficiencies with cheap tokens today will find agentic scale elusive tomorrow.” (Will Sommer, Gartner, March 2026)

And this is where the real bet lies. When intelligence becomes commoditised, and Gartner says it will, the advantage will no longer be in the model, which will have become a commodity, but in what is built around it: the harness, the trace, the business precedent. Model sovereignty is not an ideological comfort. It is the condition for architecting, from today, on the swappability of the brain. Open, stable models whose properties we know are not an end. They are the raw material of an intelligence we genuinely control. The key is not the price of the model. It is the ability to build on top of any model.

Barrier 2 | Reproducibility

A generic agent applied twice to the same contract does not always produce the same result.

This is not a configuration error. It is a structural property of LLMs, stemming from the very way they execute on hardware. It is not a bug you fix.

It is a major obstacle for an underwriting audit system: how do you present a decision to a broker, a client, or a regulator when the system itself could not reproduce it?

Reproducibility must therefore be achieved above the model, through the architecture that encapsulates it. The real work happens on the harness: the deterministic execution wrapper that surrounds the model.

The key insight opposes output to outcome. The model produces an output that varies from one run to the next. But there are infinitely many outputs, regardless of form, that produce the same outcome for the same inputs. The harness does not seek to freeze the output; it guarantees the outcome. Reproducibility becomes a structural property of the system, not a coincidence of the model.

And the true objective is richer than it appears: capture the depth of the client’s rules, expressed in natural language, and convert them into precise, reproducible, audited outcomes. The LLM generates the logic; the harness constrains it to a deterministic result. Every decision can be linked to a named, dated, re-queryable version of knowledge; every execution remains reproducible, even as the rules evolve. This also caters for a model you control: the harness only guarantees the outcome if the brain it encapsulates remains stable, and a model withdrawn or modified without notice undermines the guarantee itself. Hence, once again, the bet is on open models.

Barrier 3 | Traceability and the regulatory framework

The European regulatory framework (the AI Act, and supervisors’ expectations under Solvency II, the IDD, and DORA) is now clear on AI in insurance: record-keeping, explainability, human oversight, a clear audit trail.

AI providers sometimes underestimate th fact that the insurer remains responsible for the AI system it deploys, whether it built it or not. Responsibility cannot be delegated to the model.

An agent that cannot cite its external sources, explain its reasoning, and produce an auditable trail is not enforceable. Not against the regulator. Not against the reinsurer. Not against the policyholder in the event of a dispute.

Traceability is not a nice-to-have. It is the condition for AI’s admissibility in underwriting. And those who build it by design, rather than bolting it on afterwards, are the only ones who can truly deploy. Every alert produced must be traceable to a chain of identifiable inputs: an underwriting guide PDF, external data with its source, a validated transcript, a versioned business rule, structured policy data. Every alert must state why there is cause for concern, on what evidence, and with what level of confidence. This is not a reporting function added as an afterthought but a compulsory comoponent of the architecture.

The framing that makes agents work

This is precisely the framing that makes agents work in constrained domains: clear boundaries, measurable outcomes, obvious guardrails. In cybersecurity, Sophos MDR handles 52% of cases end-to-end with AI and no human intervention, at a median response time of 89 seconds. We’re talking about complex, sensitive tasks, yet operational because bounded by expert infrastructure.

This framing is all the more necessary in commercial P&C underwriting: domain boundaries are blurred, and results are not immediately measurable. And that is where one final piece is needed, one that no generic model and no generic architecture carries.

Business knowledge, or what the architecture does not contain

Architecture is necessary, but it is not sufficient.

What fuels it is production business knowledge: the underwriting precedent. The gap between the textual underwriting guide and what underwriters actually do. This gap cannot be demonstrated; it is acquired in production, policy after policy, alert after alert. The key lies in those seven years of accumulated portfolio analysis at Continuity; one that no generic LLM carries, or no generic architecture captures.

Conclusion

Generative AI does not replace underwriting engineering. It demands it.

Underwriting is being rebuilt around AI. The winners will not be those with the best model. They will be those who have built the system capable of turning this commoditised brain into a trustworthy, reproducible, auditable underwriting intelligence that strengthens with every cycle, fuelled by years of production knowledge that only the profession holds.

Antoine, Chief Science Officer, Continuity