AI in insurance, checked dailyThursday 17 September 2026
News, findings and tests. Every item with its source, its evidence and what it means for a book of business.For agencies, MGAs and carriers
GuideAuthored editorial guidance

Count missing fields before you quote an accuracy rate

For
Agency · MGA
Evidence
Authored editorial guidance what this label means
Sources reviewed
2026-09-17
Review cycle
Every 90 days, or sooner when a source changes
Published
2026-09-17T14:18-05:00

In short

An extraction percentage is difficult to interpret without knowing what was counted. The practical question is whether the denominator includes everything the workflow needed, not only the values the tool chose to return.

NIST recommends evaluating outputs against known ground truth and documenting measurement limits; it does not prescribe an insurance-specific denominator (NIST Generative AI Profile). The definitions below are proposed editorial measures for a bounded test.

Freeze the expected fields

Before running the tool, list the fields required for each packet. Assign an expected value and source location, or an explicit status such as absent from source, ambiguous, or not applicable. Do not force a guessed answer into the answer key.

Define permitted normalization in advance. For example, decide whether equivalent date formats are acceptable and whether a monetary value must preserve its currency and basis. A technically matching number in the wrong coverage row should not pass.

Keep four measures separate

  • Expected-field accuracy: Correct field outputs divided by all scorable expected fields. An omission of an available required value counts as incorrect.
  • Omission rate: Missing available required values divided by all scorable expected fields.
  • Unsupported values: Count generated or inferred values lacking support in the supplied sources. Report these separately, including extra output fields.
  • Packet all-correct rate: Packets with every required scorable field correct and no unsupported extra values divided by all evaluated packets.

Report ambiguous and unscorable items separately, with reasons. If you measure correct “not found” responses, show that category explicitly rather than combining it invisibly with extracted-value accuracy.

A worked example, not a benchmark

Suppose a synthetic exercise contains 100 required, scorable fields. The tool returns 92 correctly, gives five incorrect values, and omits three. Expected-field accuracy is 92/100, or 92%; omission rate is 3/100, or 3%. Counting only the 97 returned values would give about 94.8%, answering a different question.

Suppose those fields belong to ten packets and only four packets are entirely correct. Packet all-correct rate is 4/10, or 40%. These figures are invented to explain the arithmetic, not observations about any product.

Do not average away the consequential mistake

Define critical defects before the test. A proposed critical class might include a wrong insured, a materially wrong limit, or a wrong policy period. Report the count and affected packets even if the aggregate accuracy appears strong.

For an adoption decision, connect these measures to review effort and the permitted use. A high field score does not by itself justify removing a human release check.

Next: Use the pilot worksheet and scoring definitions and read why confidence is not a release decision.