Put four questions in writing before any demo, and keep the answers in the file: What does the accuracy figure count, and what is the denominator? Are human corrections included in it? Is the number from production or from testing? What baseline is it compared against, and when was that measured?
A good answer has a named customer and a date, a number with a definition beside it, and a failure they fixed. “We don't have that yet,” said plainly, is also a good answer. A tool that cannot describe its own failures will describe yours to your client.
Our own tests of these claims are scheduled on the test bench: industry-code assignment on 300 single-location submissions with adjudicated ground truth, and document extraction against a frozen field schema. A test on our dataset measures performance on our dataset; it does not validate or refute a vendor's historical customer result unless the conditions match.