A Cigna medical director sat in front of a queue of coverage requests. The system had already decided each one. His job was to add a signature.
“We literally click and submit,” a former Cigna doctor told ProPublica in March 2023. “It takes all of 10 seconds to do 50 at a time.”
Patrick Rucker, Maya Miller and David Armstrong went through the internal records. Over two months, Cigna doctors rejected more than 300,000 requests through a system called PxDx. One medical director, Dr Cheryl Dopke, handled 121,000 of them. The average time spent on a case was 1.2 seconds.
That is what an AI deployment looks like when it reaches full production in insurance. It was not a pilot. It did not stall in a proof of concept. It ran at enormous scale, cheaply, for years.
Keep that in mind while reading the rest of this, because the industry has decided its AI problem is that too little gets to production.
Everyone agrees the pilots are stuck. They disagree about why
The numbers are consistent across sources that had no reason to coordinate.
Camunda surveyed 1,150 IT leaders for its State of Agentic Orchestration and Automation 2026, published in January. Sixty-nine percent said they use AI agents. Only 11% of agentic use cases had reached production in the previous year. In insurance specifically, 65% admitted a gap between their agentic AI vision and what actually exists.
Sedgwick, in March 2026, put 82% of carriers using some form of AI and 7% at what it called scalable success. Bain surveyed 160 global insurers and found 78% had adopted generative AI while 4% had scaled it meaningfully in claims. Gartner predicted in June 2025 that more than 40% of agentic AI projects would be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls, and estimated that only about 130 of the thousands of vendors claiming agentic capability actually had it.
So the diagnosis is settled. The cause is not.
Sedgwick's answer is technical. Legacy systems are the primary constraint, some of them still running COBOL, unable to exchange data with anything cloud-native in real time. On that reading, the pilots are fine and the plumbing is the problem.
Willem Paling disagrees, and he has run the experiment. As Executive Manager, Analytics and AI at IAG, Australia's largest general insurer, he watched the same stall and concluded the leap from experiment to operation is organisational rather than technical.
“We stopped doing experiments, and we focused on delivery,” he told Alchemy Crew. “We put the same delivery expectations onto our data scientists and machine learning engineers that we have for software engineers.”
IAG then shipped more AI models in two years than in the previous six combined. Same legacy estate. Different expectations.
Paling also has a warning that sits oddly beside the productivity talk: “The most underestimated risk is AI on the other side, AI attacking the evidence layer of insurance.”
Zurich went live in eleven countries
The most-cited insurance rollout of the last year is Zurich's, with Cytora, moving unstructured submission documents into structured risk data for commercial underwriting.
Manual triage time went from 75 minutes to 15. Seven countries and two lines of business by late 2025, four more countries and four more lines in the first quarter of 2026, then a self-serve model on twelve-week deployment cycles.
“We're now at 11, and then we go to 22 by the end of the year,” said John Martin, an Application Portfolio Manager at Zurich.
The method is the interesting part, and it is unglamorous. Monthly go-live drops inside three-month delivery cycles. An Intake Master Schema of reusable core fields shared across every line of business, so country eleven is a configuration exercise rather than a new build. Native multi-language output, which removed the custom local builds that usually eat a rollout. Local underwriter engagement to drive adoption, because a tool underwriters route around is not in production in any meaningful sense.
Two caveats. This account comes from the vendor's own write-up, and the accuracy claim moves depending on where you read it: Cytora's page says 95% or better, some trade coverage says 98%. A five-year commitment to a platform deserves a number that holds still.
The ten-month result and the ten-year one
Dr Magdalena Ramada Sarasola is Senior Director and Global Insurtech Innovation Leader at WTW's Insurance Consulting and Technology division. Speaking to Insurance Business in August 2026, she cited one insurer's work producing over $13m in annual savings on claims costs and 95% first notice of loss decision accuracy within ten months.
Her explanation of why most attempts do not get there is the sharpest thing said on this subject all year.
“In practice, orchestration and contextualization beat raw model intelligence every time.”
The hard part, she argues, is not the model. It is that an experienced underwriter or claims handler “carries decades of pattern recognition, market intuition and contextual judgment” that was never written down, that they cannot fully articulate, and that shows up differently on every non-standard case. A pilot that automates the documented process automates the easy half and then discovers the other half was load-bearing.
Bryce Engelland of the Thomson Reuters Institute put the same problem in plainer terms: “If your firm runs on institutional memory, workarounds, and a kind of just ask Linda problem-solving process, then the system will eventually break down.”
Which brings up Lemonade, usually held out as the example of AI-native speed. On its second-quarter 2026 earnings call, president Shai Wininger said moving from evaluation to implementation “can happen in a matter of hours”, and reported a best-ever loss adjustment expense ratio of 5%. Chief executive Daniel Schreiber was blunter about where that came from: “These 10 years of hard work at building the technology that we've built is throwing off results.”
Hours to deploy, after a decade of building the thing you deploy into. The number that gets quoted is the first one.
A court in Minnesota asked for the paperwork
In March 2026, a magistrate judge in the District of Minnesota ruled on a discovery dispute in Estate of Gene B. Lokken et al. v. UnitedHealth Group, a case brought in November 2023 over Medicare Advantage denials of post-acute care. At issue was nH Predict, a tool from UnitedHealth's naviHealth subsidiary built on a database of roughly six million patients, which estimated how long a patient with a given profile should need rehabilitation. The plaintiffs largely won. The court ordered production of documents analysing and discussing the tool, materials on the naviHealth acquisition and the projected cost savings behind it, records on what the tool was designed to do and whether it was meant to displace a physician's judgment, and the identities of the people who trained staff to use it. It also rejected the argument that pre-2019 records were irrelevant. What the court declined to order is the part worth sitting with: it kept the source code, the algorithm's rules and the underlying data out of reach, and refused broad discovery into revenue and profit.
So the discoverable artefact was never the model. It was the record around it: what it was for, who decided, what savings were promised, how staff were told to use it.
That record is exactly what a pilot does not produce.
Europe made the same demand prospectively. Under the EU AI Act, AI systems that assess risk and price life and health cover for individuals are high-risk under Annex III, and those obligations took effect on 2 August 2026, eighteen days before this was written. Risk management across the lifecycle, data governance, technical documentation, logging and traceability of pricing-relevant outputs, transparency for deployers, and human oversight with real authority to intervene. Article 26 puts duties on the deployer, not only the vendor. Non-compliance runs to €15m or 3% of global turnover.
Ramada Sarasola's shortest quote is the one that matters here: “Regulated decisions need human accountability.”
Cigna's PxDx had a human. He had 1.2 seconds.
What the deployments that survived have in common
Read Zurich, IAG, WTW's ten-month case and Lemonade together and the pattern is not about models at all.
A reusable schema, so the second line of business is configuration rather than a rebuild. Delivery discipline borrowed from software engineering, with dates and owners. Rules held where a business user can change them, because the first version of any appetite rule is wrong and the correction cannot wait for a release cycle. Human gates placed where a regulator will ask, with enough time behind them to be real. And a log of what the system decided and why, generated as a by-product of running rather than assembled later under a discovery order.
None of that is exotic. It is ordinary software discipline applied to workflow, rating and claims, and it is unglamorous enough that pilots skip it and then cannot scale without it. Openkoda's bet is that these are configuration you own rather than a vendor roadmap you wait on: a product builder a business user can actually change, rating tables with effective dating, integrations that carry the data in, and an audit trail that exists because the platform writes one. The market for insuring AI is being built on the same assumption, which is that the record of a decision is the product.
Kurt Petersen of Camunda, summarising 1,150 responses: “Trust remains the key barrier to adoption. Without clear guardrails and visibility, agents will stay at the edge.”
Staying at the edge is the failure mode everyone is measuring. It is not the expensive one. The expensive one is a system that reaches production, runs for years, works exactly as designed, and leaves 1.2 seconds of human judgment on a decision that ends someone's care.