firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

When impressive output fails the real-world test

Technology buyers are trained to admire visible capability: polished answers, exhaustive reports and features that look impressive in a demonstration. Firmulate’s latest management wargame offers a useful warning. The participant that appeared most diligent also delivered the weakest overall result.

Opus 4.8 produced the deepest analyses and added 80 learned rules to its playbook, more than any other participant. Yet it finished last with a score of 73. Its failure was not a lack of intelligence or awareness. It identified the crises, resisted manipulation and did much of the work needed to win a valuable customer. Then it failed to complete the close.

That makes Opus 4.8 less a cautionary tale about a bad model than a character study of an extremely conscientious worker whose effort did not reliably become impact. For companies evaluating AI agents, that distinction matters more than another dazzling response in a chat window.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A brutal week with nowhere to hide

Firmulate placed frontier models in charge of the same small software company during its worst week. Each faced identical customers, crises and temptations. Decisions were versioned and auditable, allowing observers to compare management behavior rather than isolated answers.

The simulated company was not an easy assignment. It had 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown was public, every workday was versioned, and its evolving playbook contained more than 680 self-learned rules. The experiment is live and watchable, so the published results are a continuing record rather than a staged product vignette.

In the final July 2026 Crucible League, gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. A single breach of trust, however, capped the total under Firmulate’s principle that “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are available on the Firmulate benchmark page.

The missing signature

The central result was strikingly mundane. All models spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap plainly: “Same diagnosis, same pitch — no signature.”

The crucial commercial fact was not sitting conveniently inside the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that followed the trail found it, used it and won the deal at full price, worth +€4,583 in monthly recurring revenue.

Opus 4.8’s thoroughness should have made it well suited to that kind of investigation. Its problem was not an inability to reason deeply. It was prioritization and follow-through. The model generated more accumulated guidance than its peers, but the close remained on the table. Elsewhere, its discipline slipped when it repeatedly attempted to write into a locked department instead of escalating the blockage.

That is a recognizable workplace failure. A manager can produce excellent analysis, document every lesson and still underperform if the decisive next action is delayed or abandoned. More process may even make the failure harder to notice because the activity itself looks productive.

Strong judgment under pressure

A fair assessment must also recognize what Opus 4.8 and the rest of the field did well. The social-engineering tests included fake chief executive messages escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused the attempts.

Kimi K3 captured the appropriate posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Its performance deserves one methodological qualification: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

Opus 4.8 therefore did not lose because it was reckless or easily deceived. It preserved trust while showing impressive analytical depth. It lost because safe, thoughtful work was not converted into the highest-value outcome. Firmulate also found the same weakness, in milder form, across all four models in the original experiment. The lesson is broader than one profile.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Measure completion, not just cognition

The gadget industry often sells intelligence through smooth demonstrations. Business agents require a harsher standard. They must read the relevant files, recognize what matters, act on the evidence and finish what they start without surrendering judgment under pressure.

Firmulate’s 242 real, unedited management decisions also power a model-guessing quiz, underscoring how difficult it can be to identify systems by prose alone. The more revealing differences emerge through sustained behavior: which agent escalates a blockage, protects trust and secures the signature.

Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. That is the practical implication of Opus 4.8’s result. Diligence remains valuable, but volume is not execution. The best AI worker is not necessarily the one that writes the most rules. It is the one that identifies the decisive task and carries it across the line.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI follow-up automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI productivity enhancement software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Ultimate Guide To Partnering With CodeAI For AI Advancement

OpenAI has announced a partnership with CodeAI aimed at preparing the first AI generation, focusing on AI literacy and education, though details remain undisclosed.

10 Best AI-Powered Student Planners For Smarter Study Schedules In 2026

Discover the 10 best AI-compatible student planners for smarter study schedules in 2026, highlighting features, AI integration, and suitability for students.

Vertigo relief app

A new vertigo relief app aims to assist adults with BPPV using guided maneuvers and head-tracking, potentially transforming at-home vertigo management.