Why claims AI pilots stall before production
60% of insurers never get past proof of concept. After analysing 50,000 claim records, we can confidently say it isn't the model that's the problem.

Most insurers testing AI on claims struggle to reach production, and the industry has settled on governance and data readiness as the reason. We think that's accurate, but too vague to act on. When we analysed about 50,000 claim records from a live motor book, the problem turned out to be specific: the records don't describe the process they came from. Audit your own data before you buy anybody's model.
Capgemini’s 2026 World Property and Casualty Insurance Report found that 60% of P&C insurers remained in exploration or proof-of-concept stages. It identified roughly 10% as “intelligence trailblazers” that had scaled AI as a core operating capability, while 55% reported no clear return on AI investment.
The industry diagnosis is broader however. Capgemini points to an “architecture mismatch,” in which AI investment outpaces an insurer’s ability to redesign workflows, measure outcomes, assign ownership, and build organisational adoption. It also identifies legacy systems and poor data quality as barriers to execution.
It's also abstract enough that it’s challenging for people to act on. So, to resolve that challenge, we read a live book of motor claims end to end, to ensure that what we build will succeed.
Laying the ground work
We took over a claims operation running across a manufacturer's authorised service network. The data had a lot of interesting points.
The extract covered 49,941 claim dossiers and 32,149 claim requests across eight carriers, from May 2023 to June 2026. Monthly volume had grown from roughly 250 claims to between 4,000 and 4,700. A working book of live claims at production scale, in a market where vendors are actively selling automation into claims.
We were not auditing the carriers. We were trying to answer one question before committing to a design: if we automated a step, could we prove afterwards that it had worked? The answer turned out to have very little to do with the step itself.
But three things in that book had to be resolved or they’ll break anything we build:
- What we could not measure
- Where the written process and the real one diverged
- Which counterparties never produced the record at all
Let’s dive in.
AI’s plausible deniability problem
Of 23,300 paid claims, only 36.4% carried both timestamps you need to measure the approval step.
That's an average across the five carriers that had paid claims in the window, and we found the spread underneath it told us far more than the average did.
One carrier recorded both the quotation date and the approval date on 99.8% of its files. A second recorded both on 44.7%. A third recorded an approval date on 15.3% of files and a quotation date on none at all, across 8,996 paid claims. A fourth recorded an approval date on 99.1% and a quotation date on zero.
Let’s pause here.
A carrier may have approved a quotation without the quotation event being recorded in this extract. We read the pattern as a recording gap rather than a slow process: the field was not populated.
Train a model on that book and it learns the gap, not the quotation date.
An automation built to measure approval time could produce a reliable number only for the carrier whose files consistently contained both required timestamps.
For the others, the measure would be incomplete or impossible to calculate from this extract.
Off the rails
The defined process and the operating process are different documents. That gap isn’t small.
Opening a dossier straight to quotation, skipping inspection entirely, happened 14,404 times. Skipping both quotation and acceptance happened 11,053 times. Skipping eight consecutive steps in one move happened 7,510 times. Moving backwards from payment guarantee to compensation approval happened 7,919 times.
Seven dossiers even returned from a terminal state, which the process says is impossible.

The data does not establish the cause of those patterns. They may reflect a writable status field, cross-organisation workflow, system configuration, integration behaviour, user workarounds, or incomplete data capture.
But this will cost you the moment you try to automate on top of it.
An agent designed to require the documented sequence may stall or reject cases that follow the observed alternative patterns, unless it is explicitly designed and tested to handle them.
Without that context, users may interpret routine exceptions as platform failure because the recorded status transitions did not follow the documented sequence.
Three of eight
The last one is perhaps the most innocuous, but it's one that breaks the most.
Three of the eight carriers had no claim requests at all. Zero rows, against thousands of dossiers each. The other five ran roughly one request per dossier, between 0.55 and 1.35. We have confirmed with one of the three that its operation opens a dossier straight from the hotline. For the other two we have the pattern and not the confirmation. An absolute zero across thousands of files does not look like missing data, but we have not yet asked them.
The two-stage model used by five carriers is not universal in this extract. A platform that requires a claim-request record would not fit the observed data for three carriers, subject to confirming the workflow or integration design of the two not yet validated.
The same trap catches your training data, and unfortunately, it’s an easy part to miss. A model trained only on the five carriers that have claim-request records would underrepresent the alternative intake pattern. Unless it is designed for that pathway, it may route or interpret those cases poorly. Those three are not a rounding error either.
Between them, they account for 19,858 of 49,941 dossiers (39.8% of the book). If the proposed automation or training set assumes a claim-request record, that assumption would not hold for almost two-fifths of the dossiers in this extract.
Why the pilot did not warn you
Put those three findings together and the pilot-to-production gap explains itself.
Many pilots run on a prepared extract: complete rows, one carrier, a clean time window, and the fields selected for the use case. It is a fair test of the model but a poor test of everything the model will actually sit on.
Production runs on the book above. Two thirds of it can’t be measured on the step you wanted to automate. A large share of it moves in ways the process says cannot happen. Three counterparties in eight do not generate the record the model expects to read.

At no point in this analysis did the model appear to be the primary constraint. That is why Capgemini’s finding of 72% of P&C AI investment going to technology and infrastructure, versus 28% to change management, is significant.
The imbalance can leave organisations without the workflow redesign, adoption, ownership, and measurement needed to scale.
The reconciliation we run first
When we reconciled our own service-level engine against the commercial agreement, we found a three-way divergence. For one damage band the engine was measuring against 3.2 days where the commercial agreement said 5. Subject to confirmation of the operative agreement, definitions, scope, and any amendments, carriers in that band were being measured against a stricter threshold than that agreement specified.
Every party was worse off for it, including the ones the stricter measurement appeared to favour, because a performance figure that cannot be reconciled to a contract cannot be used for anything. It should not be used as the sole basis for a bonus, dispute position, or regulatory representation until it has been reconciled to the applicable agreement and validated.
So what we can state confidently here is that the sequence is the point:
- Measure the clock
- Reconcile the measurement to the agreement
- Only then attach money, penalties or automation to the number
Run it on your own book first
None of this requires a vendor, a budget or a pilot. Four queries against your own claims data will tell you most of what you need before anyone demonstrates anything.
- For the step you most want to automate, what share of closed claims carries every timestamp needed to measure it? Run it per counterparty, not in aggregate. The average will hide the carrier that’s recording nothing.
- How often does a claim move in a way your process documentation does not allow? Count the transitions, name the top five pairs, and check whether they are concentrated in one counterparty or present in all of them. Concentrated is a data problem. Widespread is how the work actually runs.
- Which defined states have never been used, by anyone? A state with no production instance is a branch your acceptance testing should not be spending time on, and a path no model should be trained to expect.
- Does any counterparty skip a stage entirely? Zero rows against thousands is a different finding from a low count. It points at an operation shaped differently rather than at missing data, but confirm it with the counterparty before you build on it.
If the answers are uncomfortable, that is a useful outcome. They were certainly uncomfortable for us.
Start from where the data is
The industry reading is correct as far as it goes. Pilots stall on governance and data readiness. Buying a better model will not change that.
What the numbers add is that data readiness is not a binary decision of meeting one condition or not. It's specific, and different for every counterparty you work with. You can also measure it today, with queries you already know how to write.

So the order matters more than the shortlist. Read your own book, counterparty by counterparty, before the demo rather than after it. A vendor can profile the data, but only you and your counterparties can validate whether it represents the operating process, contractual definitions, and exceptions the solution must support. Many pilots are not designed to discover those gaps.
When your foundation is strong, then building on it creates a stable ecosystem. If it’s shaky, then no top-tier model can truly solve those structural issues.
Have a question about our methods? Let’s compare notes.