The confidence gap in insurance AI is an evidence gap
Insurers already have the tools. What most cannot produce on short notice is a governed trail connecting one decision to what the system read, what it weighed, and who signed off.

Every insurance AI survey from this year tells a two-part story, and the coverage keeps reading only the first half. The first half: adoption is real. Most carriers already run an AI tool somewhere inside underwriting, and more are shipping this year than last. The second half gets less attention, and it is the half that matters. A 2026 Grant Thornton survey of underwriting leaders found only a minority felt highly confident their organization has a clear AI strategy, and a smaller share still said they could pass an independent governance review inside 90 days using evidence pulled from one place. The coverage described a confidence gap in insurance AI. That framing misses the mechanism. Confidence is a feeling a person reports about themselves. What the survey measured is whether one underwriting decision, picked at random, produces its own record. Most cannot, and no amount of reassurance changes that.
The survey names the wrong variable
The number under the headline says less about confidence than about what a governance review has to look at once someone finally asks for it. A board that approved an AI policy can produce that policy in an afternoon. The harder ask is the trail for one specific decision, on demand, with nothing hand-assembled overnight: what the model read, what it weighed, who saw the recommendation, and what a person did with it. That trail either exists as a byproduct of how the system runs, or it does not exist. Nobody rebuilds it retroactively once the case is closed.
Underwriting is not the only place this shows up. A separate wave of enterprise reporting this year has tracked the same pattern in agent pilots generally: mandates to deploy are common, production is rare, and the stated reason is almost never that the tool failed a demo. The pilot stalls at the point someone with governance authority asks a question it was never built to answer. Insurance carries the sharpest version of the problem, because an underwriting decision already answers to a board, an examiner, and a policyholder who can dispute the outcome. An examiner does not ask what tools a carrier has deployed. An examiner picks a file and asks to see the reasoning behind it, on whatever week the visit lands.
That timing detail is the whole story compressed into one sentence. A board update or a strategy memo can be scheduled for a date that suits the person presenting it. A file request cannot. The carriers who scored low on review readiness this year are not behind on tools. They are behind on the one property that cannot be produced under a deadline: a record that already existed before the deadline arrived.
What a governance review asks
A governance review does not ask whether a company has thought about AI. It asks a narrower question, over and over, one decision at a time: for this specific case, show me what happened. Architecture answers that question. A policy document cannot.
Every action a Nodes system takes carries a signed Decision Trace: what the system read, where it read it, why it weighed the case the way it did, and what input a person gave before anything executed. The trace is queryable, not narrated after the fact. A reviewer does not ask an analyst to remember which files fed a recommendation eight months ago. The reviewer pulls the record.
Before that trace closes, a person approves, edits, or declines the proposed action. That gate is a blocking control. A person can decline it, and the system does not proceed on its own while a decline sits unread. Nothing executes until a human with the authority to decide has decided.
The record does not stop being trustworthy the day the underlying model changes. Before a new version takes over any scored decision, it runs in shadow against the version it would replace, and it only takes the seat once it clears that bar. A case decided in March does not depend on last year's model being rebuilt to explain itself a year later. It depends on the trace from the model that actually decided it, kept intact and attached to that case forever.
Insurance adds one more layer where a decision crosses a line the carrier draws for itself. The bind decision is where this shows up first: the moment a submission moves from quote ready to a commitment the carrier is on the risk for. A second signer, a distinct named person, countersigns before that step executes, and the record shows who signed, when, and what they saw. A memo describes what should happen. A countersignature records what did happen.
None of this makes a review easy. It makes a review possible on the timeline a board sets, instead of the timeline a team can improvise once someone finally asks.
Confidence does not leave a trace
The instinct after a survey like this is to fix confidence directly: more training for the underwriting team, an update from the board, a dashboard showing adoption climbing. Each of those is real work, and none of it produces the artifact a reviewer asks for. A system prompt full of good intentions is not a governance model either. A system prompt describes what a system should do. A description is not evidence of what it did. Preferences drift as the underlying model changes; a record does not drift, because a record is a log of one event that already happened, not an instruction that gets rewritten next quarter.
The same confusion shows up in the speed conversation. Buyers ask what control model lets a vendor move that fast, assuming speed and oversight trade against each other. They stop trading against each other once the trace is a byproduct of the architecture rather than a document someone assembles after the fact. The carrier that can hand a reviewer one underwriting decision in an afternoon did not get there by moving slowly. It got there by building the record into the pipeline before the first case ran through it.
A confidence campaign changes how underwriting leaders describe the program in a meeting. What it produces once the meeting ends and someone asks for the file is unchanged, because a description was never the thing being asked for.
Two carriers can run the identical model on the identical case and still land on opposite sides of a review. The one whose system logged the read, the weighing, and the sign-off as the case moved through it walks into the review with an answer already assembled. The one relying on a policy binder and an analyst's memory of a case from March is reconstructing an answer under a deadline, and reconstruction is where inconsistency creeps in: a detail remembered differently, a step nobody wrote down, a version of events that does not quite match what the file shows elsewhere. The review does not fail on intent. It fails on reconstruction.
What already holds up under one
Nodes has made this same argument about a different high-stakes decision at the same kind of carrier, the hiring decision, and the mechanism does not change when the decision changes. Four years of production data, 10,765 agents, at a Fortune 500 insurance carrier: every score, every recommendation, and every human decision on top of it recorded as it happened. The deployment sits inside the carrier's own cloud, single-tenant, with no data egress, and the weights stay in the carrier's environment rather than on a Nodes server.
Contract to production took 34 days at a carrier that had already turned away six AI hiring vendors over eighteen months, every rejection on architecture. The question that closed those eighteen months was never whether the vendor felt confident. It was whether the record would exist before anyone needed it, and that is the same question an insurance governance review is asking about underwriting today. An architecture that produces a queryable trace, a blocking approval, and a second signer on the decisions that call for one does not have to be rebuilt by industry. The bind decision is a new place to point an existing mechanism.
The next survey will ask the same question in different words. Can you show us. The insurers who can are not the ones with the most reassuring board deck. They are the ones whose systems produce the file by default, before the request arrives, because the trail was built into the pipeline instead of promised in a policy. The tools were never the missing piece. The evidence was.
Sources
- Insurance Insights: 2026 AI Impact Survey Report (Grant Thornton)
- Insurers have the AI tools, but they do not have the confidence (Insurance Business)
Saad Bin Shafiq is the founder of Nodes, serving data-sensitive enterprises. Methodology: Decision Traces.