AI in insurance, checked dailyThursday 17 September 2026
News, findings and tests. Every item with its source, its evidence and what it means for a book of business.For agencies, MGAs and carriers
Test benchSchedule

Tests.

What these are
A tool, a claim, a frozen set of documents, and what came out
Rules
Ground truth before output. Omissions count as errors. Critical defects reported separately. Vendor sees it first.
Published
None yet — the first is in preparation

In preparation

Test 001. What does "in minutes" leave out of a submission packet?

Convr introduced DocData on 8 September 2026: upload a commercial submission packet, get a summary of operations, exposures and loss history "in minutes." The claim is a capability, not a number, so the test is not whether the summary is accurate. It is what the summary omits, and whether an omission is visible to the person relying on it.

Test set
Five authorized, redacted commercial packets, frozen before the run. Expected fields assigned by two reviewers, disagreements adjudicated by a third, all before anyone sees output.
Denominator
Every expected field, not only the fields the tool attempts. A blank counts as an error.
Reported
Field accuracy, document-level all-correct rate, silent omissions, unsupported values, and critical defects — anything that would change what a client believes is covered, what it costs, or who is insured — separately.
Conditions
Tool version, settings, retries, manual cleanup, document mix and scan quality, published with the result.
Not claimed
An error rate. Five packets demonstrate failure modes; they are not a sample a rate can be estimated from.

The vendor receives the packets, the field schema and the results before publication, with two questions: what does your accuracy figure count, and are human corrections included? The reply runs in full, or we note that none was received.

Queue

Industry-code assignment (Sixfold publishes both "exceeded 75%" and "exceeded 90%" accuracy for the same task; we will ask which population each measures, then run 300 single-location submissions against adjudicated codes). Document extraction against a frozen field schema (Cytora's "95% automated digitization accuracy" has no published definition of a field). A cycle-time study needs a cooperating carrier's timestamps and is not a desktop test.

Method in full on the methods page. The test format itself was designed before the first test ran, so that the format could not bend to the result. How to design your own controlled evaluation: the pilot guide.