Questions to ask before relying
A practical review for buyers, executives, and accountable users evaluating AI-assisted analysis or a system that claims to check it.
Use this without a score. The questions below are designed to surface evidence, limits, and operating behavior. There is no universal threshold that turns the answers into a safe product or document. The declared use and cost of error still govern.
The most useful procurement question is not “How accurate is your AI?”
Ask instead:
What exactly can I rely on, for which use, and how can I inspect the basis for that answer?
The following questions apply to a document-review product, an internal build, a consulting workflow, or a direct model review.
The ten questions
1. What decision or use is the check designed to support?
A serious answer names the decision, action, audience, or reliance and explains how the check changes with consequence.
A weak answer gives one universal quality score for every document.
2. What entered the system?
Ask how the source version is fixed, which formats are supported, what could not be read, and how tables, charts, footnotes, comments, formulas, and scanned pages are handled.
Source fidelity is not clerical. Missing content can invalidate everything downstream.
3. What did you check—and what did you not check?
The answer should separate in-document support from outside truth, determinable checks from model judgment, and completed work from suggested future research.
“We verify the document” is not a scope.
4. How does each finding bind to the source?
Ask to see the exact page, passage, cell, table, or visual that caused the finding. Then ask whether the explanation reflects the author's actual claim and intended use.
A finding that sounds true but cannot be bound to the source is difficult to review and easy to misuse.
5. Which results are computed and which are model judgments?
Arithmetic, schema, completion, and exact consistency can often be checked deterministically. Support, omitted assumptions, and reasoning quality usually require judgment.
A serious system marks the difference. It does not wrap every output in one confidence score.
6. How are material findings challenged?
Ask whether a separate job tries to disprove or limit each material finding, what context that challenger sees, and how disagreement is preserved.
Multiple models do not automatically provide independent review. The design of the work matters more than the provider count.
7. What happens when required work fails or does not finish?
Ask about malformed output, timeouts, missing source regions, contradictory results, tool failure, and incomplete checks.
The responsible answer includes a stop, hold, or refusal. Partial work must not quietly appear as a complete result.
8. What changes when the configuration changes?
Ask which model, prompt, tools, context, budget, retries, and scoring rules produced the evidence you are being shown. Ask how changes are tested and recorded.
A model name alone is not a reproducible system description.
9. What record does the user keep?
The record should show the source version, declared use, checks, material findings, challenges, limits, unresolved work, disposition, and configuration needed to understand the result.
A ledger or cryptographic anchor can protect the record from quiet alteration. It does not prove that the underlying claim is true.
10. Who owns the final action and the override?
Ask when the system indicates, conditions, holds, refuses, or allows work; who may override it; what reason must be recorded; and how the work is rechecked.
The accountable human does not disappear because the review is automated.
Ask for comparative evidence
Architecture explains how a system is intended to work. It does not prove that the system works better.
Ask for same-document comparisons against:
- a strong direct frontier-model review;
- a qualified senior human review;
- the organization's current review process; and
- relevant failure controls, including deliberately incomplete or contradictory inputs.
For every result, ask for the document set, inclusion rules, configuration, scoring method, denominator, adjudication process, false positives, false negatives, and known limitations.
Be cautious with a headline rate that does not name its unit. “95 percent accurate” is incomplete without saying: accurate at what, across which documents, under which configuration, compared with whom, and with what counted as a failure.
Review the product at its intervention point
A product that scans a document at a person's request has a different contract from a product that blocks agent work at a boundary or monitors a data estate on a cadence.
For the named intervention, require:
| Question | What to establish |
|---|---|
| Target | The document, handoff, data estate, or strategic commission being protected |
| Moment | When the check occurs relative to reliance |
| Force | Indicate, condition, hold, refuse, allow, or override |
| Latency | How long the responsible check may take |
| Authority | Who acts, who may override, and who is accountable |
| Record | What evidence and state survive the event |
| Recovery | How a user repairs, reruns, or escalates the work |
The least force that responsibly protects the declared use is usually the best design. A control that people cannot use under real deadlines will be routed around.
Test the result with a real document
Choose one consequential use
Name a document and a decision that matter. Avoid a polished demo chosen by the vendor.
Create a comparison
Run the current human process and a strong direct-model review. Freeze the material and evaluation rules before seeing the results.
Include a failure control
Add a document with a known contradiction, wrong denominator, missing source region, or deliberately incomplete job. Confirm the system stops or surfaces it.
Ask the accountable reader
Can they understand the findings, distinguish evidence from judgment, and make a better use decision without reconstructing the whole process?
The buyer's job is not to find a product that claims certainty. It is to find a process whose evidence and limits are strong enough for the reliance at hand.