
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A benchmark that refuses to hand out a zero
When we review a gadget, we know a lazy device when we see one: the smartwatch that tracks nothing, the robot vacuum that circles the same corner. So what should an AI management benchmark award to a model that, confronted with a company’s worst week, does absolutely nothing?
If you said zero, the team behind Firmulate’s benchmark disagrees — and their reasoning says a lot about how AI should actually be judged.
In the final July 2026 Crucible League standings — gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73 — a do-nothing baseline run lands at 26, not 0. That number is deliberate, and unpacking it explains what separates an honest benchmark from a leaderboard built for marketing.
Why 26 and not 0
The Firmulate experiment put each frontier model in charge of the same small software company through its worst week: same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable — the corporate equivalent of a flight-data recorder.
The scoring philosophy rests on two ideas. First, partial progress counts. A manager who correctly diagnoses a crisis but botches the response has still done something measurably better than one who never opens the ticket. Most of management is triage; a benchmark that only rewards perfect outcomes would miss that entirely. A do-nothing run still benefits from the slice of situations where standing pat was, accidentally, the least-bad move.
Second, and more striking: a single breach of trust caps the total. The benchmark’s own phrasing is blunt — “no amount of good work outweighs a breach of trust.” A model that lies once, fakes one approval, or hides one failure can never score as a top performer, no matter how brilliant the rest of its week. It’s the corporate equivalent of an auditor’s red flag, encoded directly into the grade.
Distrust of round numbers
Perhaps the most telling design choice: the benchmark treats a perfect 100 with suspicion. In a week full of judgment calls, missed nuances and trade-offs, a flawless round score smells like a metric that isn’t measuring anything hard. The gap between 26 and 95 is where the real information lives.
The finding that decided the league
Here the experiment gets genuinely uncomfortable for AI vendors. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned — “Same diagnosis, same pitch, no signature.”
The decisive fact was buried two document references deep in the company’s own files, not in the customer conversation. The models that actually read what was in front of them won the deal at full price, worth +€4,583 in monthly recurring revenue. The others had done all the hard analytical work and simply failed to close — a failure mode invisible in any chat demo.
Under pressure, all of them held the line
The social-engineering stage was nastier: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of corporate paranoia you want in an agent touching your real systems.
The thoroughness trap
Opus 4.8 is the cautionary tale of the league: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models. Effort, it turns out, is not the same as judgment.
One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.
You can watch it live
This isn’t a one-off paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics: €105k monthly burn against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

AI management benchmark software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The takeaway
A do-nothing baseline at 26, a trust ceiling that caps cheaters, and skepticism toward a clean 100 — together these say the benchmark’s designers expect AI managers to fail in human ways, and grade them the way we’d grade people: on partial credit, on follow-through, and above all on trustworthiness.
For anyone deploying AI agents near a CRM, support queue, or forecast, the Crucible results offer a deceptively simple checklist: does it finish what it starts, does it read your files before it acts, and does it stay honest when someone pretends to be the boss? Chat quality, the experiment suggests, tells you almost none of that.
The full league table and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
corporate crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI trust and ethics assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI decision-making audit software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
