AI in insurance, checked dailyThursday 17 September 2026
News, findings and tests. Every item with its source, its evidence and what it means for a book of business.For agencies, MGAs and carriers
Finding 004Open, claims added as they are published

What does “95% accuracy” actually count? Ten vendor claims, as written.

Question
Whether the accuracy and speed figures published by submission-AI vendors can be tested as written
Short answer
Mostly not. None of the ten publishes a denominator; two thresholds appear for the same task; one time reduction does not match its own before-and-after.
Basis
Ten claims from seven vendor pages and one trade release, fetched and quoted on 16 September 2026
Tier
Vendor statement — quoted as published; nothing here was tested by us
Scope reviewed
Six vendors, ten claims: Heron Data, Cytora, Sixfold, Federato, Convr, Roots Automation — not a market survey
Reviewed
16 September 2026

The finding

A percentage without a defined denominator cannot be evaluated from the published statement alone. Of ten quantified claims we read, the one with a stated baseline is the only one an underwriter can start to test, and even it does not say what a “field” is.

This is not an argument that the tools are bad. Several may be excellent. It is an argument that the published figures, as written, are not reproducible, and that a producer or MGA evaluating one should ask four questions before the demo rather than after the contract.

The claims

Quoted verbatim; “no denominator” means the page gives no base population
VendorClaim, as writtenMissingClass
Heron Data“2× submission capacity without increasing headcount”; “decision-ready risks land in your PAS in minutes, not hours”No definition of accuracy, no period, no nVague
Heron Data“cutting manual workload by up to 80%”“Up to”; no qualifying conditionsVague
Cytora / Arch“Clearance performance at just 44.9% within four hours — well short of an internal target of 80%”; “95% automated digitization accuracy achieved”; “90%+ clearance accuracy achieved in testing”Field undefined; 90% is “in testing”; the stated baseline is for timeliness, not accuracyConditionally testable
Cytora / Zurich“Manual triage time has dropped from 75 minutes to 15 minutes, while digitisation accuracy has increased from between 70% and 80% to 98%”“Digitisation accuracy” undefined; sample not disclosedConditionally testable
Sixfold / AXIS“exceeded 75% accuracy” and, elsewhere, “exceeded 90% accuracy” in applying industry codes; “AXIS analyzed more than 15,000 applications”Two thresholds for one task; 15,000 is usage, not an accuracy denominatorConditionally testable
Sixfold“50% to 97% faster processing, hit ratios up 15%, and 30% more GWP per underwriter”A 47-point range with no nVague
Sixfold“One carrier … reaching go/no-go decisions in 5 minutes”; “Another one cut days-to-quote from 16 to 7 days”Cohort and period not disclosedConditionally testable
Federato / Velocity Risk“from 21 days and 11 hours to 2 days and 8 hours”; “an 89% reduction”; “3.7x the percentage of our bound policies that meet our definition of high appetite”Cohort rules and timestamps needed; “high appetite” is the customer's own definitionConditionally testable
Convr / MSIG USA“reduced submission processing time to one hour”Start and end events, statistic, document mix undefinedConditionally testable
Roots Automation / Eastern Alliance“saved the insurer more than 2,700 human hours since Q1 2023”; priority mail “to one hour from five days”; “100-fold improvement”Accuracy only qualitative; the same piece cites an unsourced “40 percent operational capacity” survey we do not repeatConditionally testable

Sources: herondata.io (undated) and its ACORD action guide (12 Nov 2025); Ivans/Applied Systems, Arch Insurance case study (©2026); FinTech Global, 18 May 2026; sixfold.ai AXIS case study and feature releases (undated, observed 16 Sep 2026); federato.ai Velocity Risk case study (go-live Sep 2023); convr.com case studies (undated); Insurance Innovation Reporter, 10 Oct 2024.

Arithmetic that does not reconcile

The Arch case study reports both a “70% reduction” in intake-to-decision time and “from 2–3 days to 2 hours.” On calendar days the second implies roughly 96–97 percent; on eight-hour working days roughly 88–92. Neither is 70. Some convention is missing (cohort, clock, or metric) and the right response is to ask the vendor to reconcile, not to pick the figure that suits the argument.

Sixfold's two thresholds are not a contradiction either: more than 90 also exceeds 75. But a reader is owed the population, the stage and the definition behind each, and the page gives none.

What to do with this

Put four questions in writing before any demo, and keep the answers in the file: What does the accuracy figure count, and what is the denominator? Are human corrections included in it? Is the number from production or from testing? What baseline is it compared against, and when was that measured?

A good answer has a named customer and a date, a number with a definition beside it, and a failure they fixed. “We don't have that yet,” said plainly, is also a good answer. A tool that cannot describe its own failures will describe yours to your client.

Our own tests of these claims are scheduled on the test bench: industry-code assignment on 300 single-location submissions with adjudicated ground truth, and document extraction against a frozen field schema. A test on our dataset measures performance on our dataset; it does not validate or refute a vendor's historical customer result unless the conditions match.