What to measure
A practical evidence plan for comparative quality, failure handling, decision usefulness, and future-facing work.
Measure the promise that matters. The open question is not whether a system can produce a polished review. It is whether the process gives a responsible reader a stronger basis for the declared decision than the alternatives—and whether it stops responsibly when it cannot.
The case for decision readiness should become more specific as evidence accumulates.
That requires a scorecard with named units and denominators, not a forecast about where the market will be in eighteen months.
The evidence still owed
Comparative quality
Run the same documents through a strong direct frontier-model review, a qualified human review, the current organizational process, and the proposed integrity process. Blind adjudicators to the system identity where practical.
Recall and precision
Measure the share of material issues found and the share of reported issues that are real under the test rules. Report both. A system that catches more by overwhelming the reader with false alarms has not solved the problem.
Stability
Repeat runs under frozen conditions. Then change one material part—model, prompt, tool, budget, or source representation—and record what moves.
Decision usefulness
Test whether responsible readers understand the result, distinguish evidence from judgment, and make an appropriate use decision. Comprehension is part of the product outcome.
Repair value
Measure whether authors fix material problems faster and more completely, without silently deleting true findings or introducing new ones.
Failure handling
Include malformed, incomplete, contradictory, and unreadable inputs. Confirm that the system holds or refuses rather than presenting partial work as complete.
A minimum scorecard
| Measure | Unit and denominator to name | Why it matters |
|---|---|---|
| Material-issue recall | Real material issues found / all adjudicated material issues | Shows what the process misses |
| Finding precision | Valid material findings / all material findings reported | Shows review burden and false alarms |
| Run completion integrity | Valid complete runs / all attempted runs | Distinguishes honest refusal from false completion |
| Disposition stability | Same use decision / repeated frozen runs | Shows whether the final action is dependable |
| Reader comprehension | Readers answering defined interpretation questions correctly / readers tested | Shows whether the record can be used responsibly |
| Repair completion | Material issues resolved without new material defects / documents repaired | Shows whether findings improve the work |
| Time to responsible use decision | Elapsed time from admitted source to disposition | Tests fit with real workflow |
| Bypass rate | Decision-lane uses outside the controlled path / all observed Decision-lane uses | Shows whether the control survives deadline pressure |
Every measure needs the document set, task, configuration, inclusion rules, adjudication method, and failure rule. A percentage without those details is not decision evidence.
Scenario planning without false certainty
Scenarios, backcasting, and recommendations sit farther from source evidence than direct claims. Measure and display that distance.
For each scenario, record:
- Starting facts — what the admitted evidence directly establishes.
- Assumptions — what must be true for the branch to hold.
- Branch conditions — the events or thresholds that separate one path from another.
- Horizon — how far forward the scenario reaches.
- Disconfirming signals — what would weaken or close the branch.
- Implications — what changes if the branch occurs.
- Action and reversibility — what may be done now and how difficult it is to unwind.
A scenario is not verified because its inputs were checked. A backcast is not evidence that the desired future will occur. The source record can be strong while the future-facing structure remains conditional.
Watches should be operational, not prophetic
A useful watch names an object, signal, source, cadence, owner, expected latency, and action.
| Field | Question |
|---|---|
| Object | Which accepted analysis, assumption, decision, or risk does this watch protect? |
| Signal | What observable change matters? |
| Source | Where will the signal come from, and what are its rights and limits? |
| Cadence | How often is it checked? |
| Latency | How quickly after source availability should it surface? |
| Owner | Who decides what the signal means? |
| Action | Reconsider, research, repair, hold, or no action? |
A new signal should not silently rewrite an accepted view. It should create a visible reason to reconsider it.
What would weaken the thesis
The decision-readiness thesis should change if controlled evidence shows that:
- a strong direct-model review or current human process performs as well or better on material recall, precision, stability, and reader usefulness at acceptable cost;
- the added process increases noise or delay without improving the responsible use decision;
- users regularly bypass the control because the intervention is poorly placed or too slow;
- material limits disappear between the scan and later analysis;
- future-facing work presents more precision without better-calibrated support; or
- comparative gains disappear when the test documents, models, or adjudicators change.
That is not failure of the broader need. It is evidence that the proposed method, configuration, or intervention has not earned its claim.
Evidence maintenance
The active external research on this site is reviewed under a freshness rule beginning April 1, 2026.
For each new source, record the publication date, source type, population or task, exact result, unit and denominator, plain-language meaning, what it does not mean, and the outward claim it may support. Remove or retire a source when its task no longer supports the claim, a correction changes the meaning, or a stronger current source replaces it.
The website should become more conservative where evidence weakens and more specific where proof accumulates.
Start with one complete comparison
Choose one real document and one declared use
Use material representative of the work people actually bring. Freeze the document before evaluation.
Define materiality before reading the results
State which issues could change the use decision and how disagreements will be adjudicated.
Run the comparators
Use the current human process, a strong direct model review, and the integrity process under recorded configurations.
Test the reader, not only the engine
Confirm that the accountable user can see what changed, what remains unresolved, and what action the record permits.
A thin, complete comparison is worth more than a broad architecture with no measured decision outcome.