AI in insurance, checked dailyThursday 17 September 2026
News, findings and tests. Every item with its source, its evidence and what it means for a book of business.For agencies, MGAs and carriers
Working resourceAuthored editorial guidance

AI pilot worksheet

For
Agency · Producer · MGA
Evidence
Authored editorial guidance what this label means
Sources reviewed
2026-09-17
Review cycle
Every 90 days, or sooner when a source changes
Published
2026-09-17T14:18-05:00

In short

A useful pilot begins with a decision you can make from its results. Use this worksheet to define one workflow and its acceptance rules before looking at the tool's output.

This is an editorial template informed by NIST's emphasis on intended context, measurement, documented roles, and risk treatment; it is not a NIST certification or an insurance requirement (NIST AI RMF Core).

Define the evaluation

Download the blank pilot worksheet. Use one row per pilot version, and keep detailed packet-level measurements in your approved evaluation system rather than cramming them into one cell.

  • Task and boundary: State the input, expected output, reviewer, and next step. List actions the tool is not allowed to take.
  • Permission and environment: Record the authorized dataset, account type, permitted use, access owner, and relevant approval reference.
  • Ground truth: Identify who establishes the expected answer, from which source, before results are scored. Define how ambiguity and unscorable items are handled.
  • Sample design: Describe document types, quality variation, selection method, and exclusions. Record sample size without claiming that a small convenient sample represents the whole business.
  • Comparison: Define the existing manual process, output quality standard, and human time categories.
  • Acceptance and stop rules: State acceptable defects by severity, any prohibited outcomes, escalation, and manual fallback.
  • Decision: Record accept, revise, reject, or insufficient evidence, with an owner and reasons.

A worked scope example

Consider this invented pilot: “Prepare an internal inventory of documents in five authorized synthetic submission packets.” The tool may label and list documents; it may not decide appetite, infer absent facts, send messages, or write to a production system.

The reviewer creates the expected inventory first. An omitted document is counted as an error rather than excluded because the tool returned nothing. A packet with unreadable content is flagged for manual handling, with the handling rule fixed before scoring. Five packets can reveal failure modes; this example does not establish a sample-size standard or a reliable performance estimate.

Close the loop

Write the decision even when the answer is no. If the result is promising but the sample excludes difficult documents, approve only the next evaluation stage, not a wider use the evidence did not cover.

Read Design an AI pilot that can produce a no before running the worksheet. Pair it with measurement guidance so the acceptance rule and denominator describe the same task.