06 · What to measure

What to measure

A practical evidence plan for comparative quality, failure handling, decision usefulness, and future-facing work.

Measure the promise that matters. The open question is not whether a system can produce a polished review. It is whether the process gives a responsible reader a stronger basis for the declared decision than the alternatives—and whether it stops responsibly when it cannot.

The case for decision readiness should become more specific as evidence accumulates.

That requires a scorecard with named units and denominators, not a forecast about where the market will be in eighteen months.

The evidence still owed

Comparative quality

Run the same documents through a strong direct frontier-model review, a qualified human review, the current organizational process, and the proposed integrity process. Blind adjudicators to the system identity where practical.

Recall and precision

Measure the share of material issues found and the share of reported issues that are real under the test rules. Report both. A system that catches more by overwhelming the reader with false alarms has not solved the problem.

Stability

Repeat runs under frozen conditions. Then change one material part—model, prompt, tool, budget, or source representation—and record what moves.

Decision usefulness

Test whether responsible readers understand the result, distinguish evidence from judgment, and make an appropriate use decision. Comprehension is part of the product outcome.

Repair value

Measure whether authors fix material problems faster and more completely, without silently deleting true findings or introducing new ones.

Failure handling

Include malformed, incomplete, contradictory, and unreadable inputs. Confirm that the system holds or refuses rather than presenting partial work as complete.

A minimum scorecard

MeasureUnit and denominator to nameWhy it matters
Material-issue recallReal material issues found / all adjudicated material issuesShows what the process misses
Finding precisionValid material findings / all material findings reportedShows review burden and false alarms
Run completion integrityValid complete runs / all attempted runsDistinguishes honest refusal from false completion
Disposition stabilitySame use decision / repeated frozen runsShows whether the final action is dependable
Reader comprehensionReaders answering defined interpretation questions correctly / readers testedShows whether the record can be used responsibly
Repair completionMaterial issues resolved without new material defects / documents repairedShows whether findings improve the work
Time to responsible use decisionElapsed time from admitted source to dispositionTests fit with real workflow
Bypass rateDecision-lane uses outside the controlled path / all observed Decision-lane usesShows whether the control survives deadline pressure

Every measure needs the document set, task, configuration, inclusion rules, adjudication method, and failure rule. A percentage without those details is not decision evidence.

Scenario planning without false certainty

Scenarios, backcasting, and recommendations sit farther from source evidence than direct claims. Measure and display that distance.

For each scenario, record:

  1. Starting facts — what the admitted evidence directly establishes.
  2. Assumptions — what must be true for the branch to hold.
  3. Branch conditions — the events or thresholds that separate one path from another.
  4. Horizon — how far forward the scenario reaches.
  5. Disconfirming signals — what would weaken or close the branch.
  6. Implications — what changes if the branch occurs.
  7. Action and reversibility — what may be done now and how difficult it is to unwind.

A scenario is not verified because its inputs were checked. A backcast is not evidence that the desired future will occur. The source record can be strong while the future-facing structure remains conditional.

Watches should be operational, not prophetic

A useful watch names an object, signal, source, cadence, owner, expected latency, and action.

FieldQuestion
ObjectWhich accepted analysis, assumption, decision, or risk does this watch protect?
SignalWhat observable change matters?
SourceWhere will the signal come from, and what are its rights and limits?
CadenceHow often is it checked?
LatencyHow quickly after source availability should it surface?
OwnerWho decides what the signal means?
ActionReconsider, research, repair, hold, or no action?

A new signal should not silently rewrite an accepted view. It should create a visible reason to reconsider it.

What would weaken the thesis

The decision-readiness thesis should change if controlled evidence shows that:

  • a strong direct-model review or current human process performs as well or better on material recall, precision, stability, and reader usefulness at acceptable cost;
  • the added process increases noise or delay without improving the responsible use decision;
  • users regularly bypass the control because the intervention is poorly placed or too slow;
  • material limits disappear between the scan and later analysis;
  • future-facing work presents more precision without better-calibrated support; or
  • comparative gains disappear when the test documents, models, or adjudicators change.

That is not failure of the broader need. It is evidence that the proposed method, configuration, or intervention has not earned its claim.

Evidence maintenance

The active external research on this site is reviewed under a freshness rule beginning April 1, 2026.

For each new source, record the publication date, source type, population or task, exact result, unit and denominator, plain-language meaning, what it does not mean, and the outward claim it may support. Remove or retire a source when its task no longer supports the claim, a correction changes the meaning, or a stronger current source replaces it.

The website should become more conservative where evidence weakens and more specific where proof accumulates.

Start with one complete comparison

01

Choose one real document and one declared use

Use material representative of the work people actually bring. Freeze the document before evaluation.

02

Define materiality before reading the results

State which issues could change the use decision and how disagreements will be adjudicated.

03

Run the comparators

Use the current human process, a strong direct model review, and the integrity process under recorded configurations.

04

Test the reader, not only the engine

Confirm that the accountable user can see what changed, what remains unresolved, and what action the record permits.

A thin, complete comparison is worth more than a broad architecture with no measured decision outcome.