🔍 Read the full analysis: AI And The New Cost Of Confidence: Reviewing What Gets Made on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
A source report describes AI-generated work growing across mathematics, software and contract workflows while human review remains limited. The figures it cites suggest that producing drafts and results is becoming faster than checking their correctness, relevance and readiness for use.
A report published this week argues that AI is making work cheaper to produce faster than people can verify it, citing a batch of 722 mathematical manuscripts and studies of software teams and contract workflows. The development matters because review capacity can determine whether AI-generated output is safe, correct and useful—not simply how much of it can be made.
According to the source, OpenAI posed about 4,000 mathematical problems to a model and published 722 manuscripts grouped into 372 families. The average result reportedly required about three hours of compute. Some results have been formally checked using Lean, a proof assistant, while OpenAI cautioned that unformalized work could contain issues. The source contrasts this volume with the careful verification of an earlier counterexample to an Erdős conjecture by five prominent mathematicians. That example illustrates the difference between generating a candidate result and establishing whether it is sound and significant.
The report also cites software-industry measurements. Faros AI found teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, in an analysis of 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study cited by the report found 61% of AI-agent pull requests received no human review before being merged or closed.
In contracts, the source describes OpenAI’s partnership with contract-software company Ironclad and an evaluation of GPT-6 Astra on 11 tasks. The model met 55% of evaluation criteria on average, which the source characterizes as a substantial improvement over its predecessor. The remaining criteria still require attention before a draft can be relied on. The report’s figures come from different studies and settings, so they should not be treated as one unified measure of AI performance.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the Pace
The immediate consequence is that more generated work does not automatically mean more usable work. When review teams cannot keep pace, organizations may face longer queues, approve work with limited scrutiny, or reject machine-generated drafts more readily. Each response carries a cost: delays, missed defects or useful output left unused.
The report also points to a workforce issue. Experienced reviewers generally build judgment through years of doing the underlying work: writing code, drafting contracts or producing research. If AI replaces too much entry-level practice, the future pool of people able to assess its output may shrink even as demand for assessment grows. The source frames this as a risk, not a demonstrated outcome across every profession.
That imbalance could increase the value of people who can make and defend a reliable judgment. Senior engineers, specialist lawyers, auditors and scientific referees may become constraints on how much AI-assisted work an organization can responsibly put to use. The scale of any such effect remains uncertain, but review time is already a practical operational concern in the cited software data.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Bottleneck
The report brings together mathematics, software and contract work because each separates production from adjudication. A formal checker can test whether a proof follows from its stated assumptions, for example, but human specialists still have to judge whether the theorem is relevant and whether the assumptions address the intended question. Likewise, software tests only check what the tests cover; they do not establish that the requirements were right.
The comparison has limits. Mathematical manuscripts, pull requests and contract tasks are different kinds of work, and their measurements use different methods. Faros AI and LinearB sell software-related products, a commercial interest the source says readers should keep in mind when interpreting their findings. The cited peer-reviewed study is a separate source, but the report does not provide enough detail here to compare its methods directly with the industry analyses.
The report’s broader point is that technical checking can help without removing human responsibility. A person or institution may still have to approve a contract, sign off on an engineering design or answer for published research. A model can generate material, but it does not take on those duties.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The source material does not identify the full publication details, dates or methods for every cited study, and the measures are not directly comparable. It also does not establish how representative the reported results are of all mathematical research, software teams or legal work. The industry figures come in part from companies that sell code-review tools, which may shape how their findings are presented.
It remains unclear how many of the 722 manuscripts were independently reviewed, how many results were later confirmed or corrected, and how the contract evaluation’s criteria map to real-world legal outcomes. Nor do the cited figures show whether review delays eventually fall as organizations adapt their processes. The claim that AI could weaken training pathways for future experts is a plausible concern raised by the report, not a measured long-term effect in the evidence presented.
As an affiliate, we earn on qualifying purchases.
Track Validation and Training
The next useful evidence will be independent follow-up on the mathematical manuscripts, including which results receive formal verification and expert review. For software, organizations will need to track not only how many pull requests AI helps produce, but also review coverage, defect rates and time to approval. Contract evaluations will need to show how models perform against clearly stated requirements in real workflows, beyond a single average score.
For employers and professional bodies, the question is whether junior staff still get enough hands-on work to develop judgment while AI handles more drafting and production. The source offers no specific policy or timeline. For now, its findings make review capacity, accountability and training the key measures to watch alongside output and speed.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the report’s main finding?
It argues that AI output is scaling faster than human verification in mathematics, software and contract work. The evidence cited varies by field and does not establish one universal rate.
Did OpenAI publish 722 verified mathematical results?
The source says OpenAI published 722 manuscripts from work on about 4,000 problems. Some were formally checked in Lean, but OpenAI cautioned that unformalized results could have issues; the source does not say all 722 were independently verified.
What did the software studies report?
Faros AI reported more pull requests merged alongside longer review time during high-AI-adoption periods. LinearB reported longer waits before review began for AI-generated changes, while a peer-reviewed 2026 study cited by the report found many AI-agent pull requests lacked human review. These are distinct studies with different methods.
Does the report show that AI review tools cannot check AI work?
No. The report says automated checks can help, but they may not establish whether the original requirements were right, whether a result matters, or who is accountable. It does not claim that every AI-generated result requires the same kind of human review.
What remains unknown?
The evidence presented does not establish how representative the findings are, how many mathematical manuscripts will be independently confirmed, or whether review bottlenecks will ease as workflows change. Long-term effects on training future experts also remain unmeasured.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
