Design an AI pilot that can produce a no
- For
- Agency · MGA
- Evidence
- Authored editorial guidance — what this label means
- Sources reviewed
- 2026-09-17
- Review cycle
- Every 90 days, or sooner when a source changes
- Published
- 2026-09-17T14:18-05:00
A pilot should make it possible to decline the tool, restrict its use, or postpone adoption. If the only planned outcome is a success story, the exercise is a demonstration rather than an evaluation.
NIST recommends documenting intended context, testing under conditions similar to deployment, and setting minimum performance or assurance thresholds for go/no-go decisions (AI RMF Core, Generative AI Profile). The workflow below is an editorial method for putting those ideas into practice.
Choose one task and define its beginning and end. “Process a submission” is too broad if one evaluator stops after extraction and another stops after the underwriter has corrected the record. For a first test, use an explicit endpoint such as “a reviewed document inventory ready for the underwriter.”
Specify the allowed document types, business lines, user roles, and actions. Identify the system version, settings, and any human or vendor assistance. Keep unsupported cases visible rather than quietly removing them.
Select authorized examples that reflect the intended work, including difficult cases. Preserve the input version and define the expected answers before reviewing tool output. Where reviewers disagree, record the disagreement and resolve it or mark the field unscorable.
Separate configuration examples from evaluation examples. If a vendor tunes the system on a packet, do not present that packet as an untouched test. Record any later changes so the comparison remains interpretable.
Capture the manual baseline and the AI-assisted process using comparable start and stop events. Include preparation, waiting, review, correction, and rework. Track critical defects separately from the aggregate score.
For a small exploratory test, report what happened in the observed cases and describe the document mix. Do not extrapolate a handful of successful packets into a reliable error rate for an entire book; NIST specifically cautions against extrapolating from narrow or anecdotal assessments (NIST Generative AI Profile).
Proposed stop conditions include unauthorized disclosure, an unapproved external action, or repeated unsupported values in material fields. These are editorial examples, not an exhaustive incident policy. The organization should define its own thresholds and escalation owners before testing.
At the end, choose among adopt for the tested scope, adopt with restrictions, extend the test, or decline. Document what remains unknown. A tool can be helpful for one document family and unsuitable for another.
Next: Use the pilot worksheet and read how to count accuracy. No pilot results are claimed in this guide.
Make the next AI decision with better evidence.
Join the weekly Renewal Report for source-backed developments, practical checks, and clear limits on what is known.
Get the weekly reportOne email a week. Unsubscribe at any time.