Current evidence
What current 2026 research and primary records establish about AI-assisted work—and what they do not establish about any product.
Evidence policy. The active research case on this site uses work published or materially updated on or after April 1, 2026. Older studies can remain useful historically, but they do not support the current case here. Primary court records are treated separately from research.
The evidence does not supply one universal AI error rate. It supports a more useful set of bounded conclusions:
- AI use is already common in professional work.
- Review and correction remain material parts of that work.
- A coarse yes-or-no evaluation can hide partial support.
- A correct endpoint can hide weak earlier reasoning.
- Long, delegated workflows introduce failure patterns that short tasks do not show.
- A model name, citation, or completed run is not a complete account of system quality.
External research supports the problem. It does not prove that any particular review system outperforms a strong direct model review or a qualified human reviewer.
AI is already in the work
Organization adoption
Stanford's 2026 AI Index reports that 88 percent of surveyed organizations used AI in at least one business function in 2025, and 70 percent used generative AI in at least one function. This is adoption context, not a quality measure.
Professional use
The Thomson Reuters Future of Professionals Report 2026 surveyed 1,816 professionals across 62 countries. Seventy-four percent reported using AI at work at least several times a week. The report also found a gap between expected client demand for AI-enabled quality and reported delivery by professional-service providers.
Rework and explainability
The commercial Work AI Index 2026 surveyed 6,000 full-time digital workers in the United States, United Kingdom, and Australia. Forty-one percent said they had delivered AI-assisted work they could not fully explain; 77 percent reported correcting or redoing AI work in the prior month. The data is self-reported and the sample skews toward AI-heavy digital work.
These results establish scale and review pressure. They do not show how often an important document fails its intended decision.
Supported is not the same as partly supported
RAND evaluated models against 240 claims drawn from 14 technical policy reports and reviewed by 16 subject-matter experts.
In the preliminary baseline, three tested models achieved 48, 54, and 53 percent exact accuracy when classifying claims into six support categories. When the task was reduced to a binary label, the corresponding results rose to 75, 80, and 79 percent.
The important point is not a universal performance rate. It is the difference created by the rubric. A yes-or-no check can look much stronger while losing the distinction between full and partial support.
Safe reading: A document check should preserve degrees of support when those degrees matter to the decision. The RAND task concerned technical policy reports under a specific preliminary benchmark; its rates do not transfer automatically to board papers, investment memos, or another review system.
A correct endpoint can hide a weak route
A peer-reviewed April 2026 study in JAMA Network Open tested 21 off-the-shelf language models across 29 stepwise clinical vignettes and five parts of clinical reasoning.
The models performed more strongly on final diagnosis than on building a differential diagnosis and navigating uncertainty. The study's multidimensional score separated performance more clearly than raw final-answer accuracy.
The bounded lesson is useful outside medicine: a strong final answer can hide weak performance in the steps that should justify it. The measured rates remain clinical and should not be imported into business-document claims.
Long workflows create different failures
DELEGATE-52: document fidelity across long editing
This April 2026 preprint tested 19 models in simulated editing workflows across 52 structured professional domains. After 20 sessions, the three tested frontier models lost about 25 points on the authors' custom document-reconstruction measure. Longer documents, more interactions, and distractor files made degradation worse.
This is a controlled simulation and a custom fidelity measure. It does not mean that AI loses 25 percent of the facts in every document.
HORIZON: inspect the path, not only the endpoint
This April 2026 preprint studied more than 3,100 trajectories across four agent domains. It was designed to locate where long-horizon systems break, not merely whether an endpoint was reached.
The safe conclusion is that connected, long tasks expose failure patterns that short task scores can miss. Trajectory-level evidence matters.
Multi-agent chains can trade one error for another
A June 2026 preprint evaluated 500 three-agent cascades across 10 knowledge domains. The chains reduced the paper's hallucination score while also showing a smaller decline in factual accuracy.
More agents were not simply better or worse. One measure improved while another weakened. A system has to name what it checks and preserve the trade-off.
Network failures can appear between reasonable agents
Microsoft Research reported an April 2026 red-team study of a live internal platform with more than 100 agents. The team observed propagation, amplification, manufactured consensus, and loss of source visibility across proxy chains.
This is an internal platform report, not a population estimate. It supports inspecting handoffs, provenance, and intervention paths rather than assuming reasonable parts create a reliable whole.
A primary-record example: the parts were real, the total did not survive
In Trinseo Europe GmbH v. Harper, Trinseo alleged misappropriation of ten trade secrets. Its damages expert valued the case on the premise that all ten had been taken and did not provide a method for valuing a smaller combination.
The jury found that four qualified as trade secrets and were misappropriated. It awarded more than $75 million. The damages award was later vacated, and the Fifth Circuit affirmed the take-nothing judgment on damages, because the evidence did not show how much of the total belonged to the four claims that survived.
The example is narrow and human: individual parts of a case can be real while the final number still lacks the analysis needed to support it. This is one legal record, not evidence of frequency.
Trinseo Europe GmbH v. Harper, Fifth Circuit
Evaluation is a system claim
Current provider guidance points in the same direction.
- OpenAI's May 2026 evaluation playbook says an evaluation should name the claim it supports and disclose the harness, tools, budgets, retries, context handling, scoring, and validity checks. This is provider guidance, not independent product evidence.
- Anthropic's April 2026 agent guidance treats the model, harness, tools, and environment as interacting parts of the system. This is provider architecture and policy guidance.
- Google's May 2026 long-running-agent pattern uses explicit state, durable checkpoints, event-driven pauses, and human approval gates. This is official developer guidance, not comparative research.
The shared lesson is modest: a model name and a final score do not fully describe the system that produced a result.
What remains unproven
The strongest missing evidence for any decision-readiness system is comparative evidence:
- Same-document tests against a strong direct frontier-model review and qualified human review.
- Calibrated recall and precision for material findings.
- Stability across repeated runs and configuration changes.
- Decision usefulness: whether responsible readers understand the result and make a better use decision.
- Repair value: whether authors fix material problems faster and more completely.
- Failure handling: whether incomplete, contradictory, or malformed work is reliably stopped and surfaced.
Until those measures exist, “substantially more decision-ready” is a standard to earn—not a measured superiority claim.