AI And The New Cost Of Confidence: Reviewing What Gets Made
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI And The New Cost Of Confidence: Reviewing What Gets Made on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source report describes AI-generated work growing across mathematics, software and contract workflows while human review remains limited. The figures it cites suggest that producing drafts and results is becoming faster than checking their correctness, relevance and readiness for use.

A report published this week argues that AI is making work cheaper to produce faster than people can verify it, citing a batch of 722 mathematical manuscripts and studies of software teams and contract workflows. The development matters because review capacity can determine whether AI-generated output is safe, correct and useful—not simply how much of it can be made.

According to the source, OpenAI posed about 4,000 mathematical problems to a model and published 722 manuscripts grouped into 372 families. The average result reportedly required about three hours of compute. Some results have been formally checked using Lean, a proof assistant, while OpenAI cautioned that unformalized work could contain issues. The source contrasts this volume with the careful verification of an earlier counterexample to an Erdős conjecture by five prominent mathematicians. That example illustrates the difference between generating a candidate result and establishing whether it is sound and significant.

The report also cites software-industry measurements. Faros AI found teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, in an analysis of 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study cited by the report found 61% of AI-agent pull requests received no human review before being merged or closed.

In contracts, the source describes OpenAI’s partnership with contract-software company Ironclad and an evaluation of GPT-6 Astra on 11 tasks. The model met 55% of evaluation criteria on average, which the source characterizes as a substantial improvement over its predecessor. The remaining criteria still require attention before a draft can be relied on. The report’s figures come from different studies and settings, so they should not be treated as one unified measure of AI performance.

At a glance
reportWhen: Published this week, according to the s…
The developmentA report argues that AI is lowering the cost of producing work across several fields while leaving verification and human accountability as constrained parts of the process.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the Pace

The immediate consequence is that more generated work does not automatically mean more usable work. When review teams cannot keep pace, organizations may face longer queues, approve work with limited scrutiny, or reject machine-generated drafts more readily. Each response carries a cost: delays, missed defects or useful output left unused.

The report also points to a workforce issue. Experienced reviewers generally build judgment through years of doing the underlying work: writing code, drafting contracts or producing research. If AI replaces too much entry-level practice, the future pool of people able to assess its output may shrink even as demand for assessment grows. The source frames this as a risk, not a demonstrated outcome across every profession.

That imbalance could increase the value of people who can make and defend a reliable judgment. Senior engineers, specialist lawyers, auditors and scientific referees may become constraints on how much AI-assisted work an organization can responsibly put to use. The scale of any such effect remains uncertain, but review time is already a practical operational concern in the cited software data.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Bottleneck

The report brings together mathematics, software and contract work because each separates production from adjudication. A formal checker can test whether a proof follows from its stated assumptions, for example, but human specialists still have to judge whether the theorem is relevant and whether the assumptions address the intended question. Likewise, software tests only check what the tests cover; they do not establish that the requirements were right.

The comparison has limits. Mathematical manuscripts, pull requests and contract tasks are different kinds of work, and their measurements use different methods. Faros AI and LinearB sell software-related products, a commercial interest the source says readers should keep in mind when interpreting their findings. The cited peer-reviewed study is a separate source, but the report does not provide enough detail here to compare its methods directly with the industry analyses.

The report’s broader point is that technical checking can help without removing human responsibility. A person or institution may still have to approve a contract, sign off on an engineering design or answer for published research. A model can generate material, but it does not take on those duties.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The source material does not identify the full publication details, dates or methods for every cited study, and the measures are not directly comparable. It also does not establish how representative the reported results are of all mathematical research, software teams or legal work. The industry figures come in part from companies that sell code-review tools, which may shape how their findings are presented.

It remains unclear how many of the 722 manuscripts were independently reviewed, how many results were later confirmed or corrected, and how the contract evaluation’s criteria map to real-world legal outcomes. Nor do the cited figures show whether review delays eventually fall as organizations adapt their processes. The claim that AI could weaken training pathways for future experts is a plausible concern raised by the report, not a measured long-term effect in the evidence presented.

Amazon

contract review AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Validation and Training

The next useful evidence will be independent follow-up on the mathematical manuscripts, including which results receive formal verification and expert review. For software, organizations will need to track not only how many pull requests AI helps produce, but also review coverage, defect rates and time to approval. Contract evaluations will need to show how models perform against clearly stated requirements in real workflows, beyond a single average score.

For employers and professional bodies, the question is whether junior staff still get enough hands-on work to develop judgment while AI handles more drafting and production. The source offers no specific policy or timeline. For now, its findings make review capacity, accountability and training the key measures to watch alongside output and speed.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the report’s main finding?

It argues that AI output is scaling faster than human verification in mathematics, software and contract work. The evidence cited varies by field and does not establish one universal rate.

Did OpenAI publish 722 verified mathematical results?

The source says OpenAI published 722 manuscripts from work on about 4,000 problems. Some were formally checked in Lean, but OpenAI cautioned that unformalized results could have issues; the source does not say all 722 were independently verified.

What did the software studies report?

Faros AI reported more pull requests merged alongside longer review time during high-AI-adoption periods. LinearB reported longer waits before review began for AI-generated changes, while a peer-reviewed 2026 study cited by the report found many AI-agent pull requests lacked human review. These are distinct studies with different methods.

Does the report show that AI review tools cannot check AI work?

No. The report says automated checks can help, but they may not establish whether the original requirements were right, whether a result matters, or who is accountable. It does not claim that every AI-generated result requires the same kind of human review.

What remains unknown?

The evidence presented does not establish how representative the findings are, how many mathematical manuscripts will be independently confirmed, or whether review bottlenecks will ease as workflows change. Long-term effects on training future experts also remain unmeasured.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

9 Best 3D Printers In 2026

Discover the nine best 3D printers in 2026, featuring top picks for performance, speed, ease of use, and value, based on expert evaluations.

Pico Space Pro Clears FCC Following Cancelled September Reveal

Pico’s upcoming XR headset, Space Pro, has passed FCC certification following its postponed September launch, with release now targeted for Q4 2026.

Innovative Antimicrobial Search Methods With Codex And ChatGPT AI Tools

University of Pennsylvania researchers use Codex, ChatGPT, and deep learning to cut antimicrobial candidate search from years to hours, advancing drug discovery.

Is Xfinity down? Thousands report TV service issues

Thousands report widespread Xfinity TV service disruptions, with reports peaking at over 50,000 searches. Service issues confirmed by the provider, ongoing investigations.