
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
A benchmark win is not a business win
Technology buyers have learned to scan coding leaderboards and chat arenas before deciding which AI model deserves attention. Those comparisons can reveal whether a model writes strong code or produces convincing answers. They say much less about what happens when an agent must operate through a churn wave, price increase, downround or PR crisis while customers, cash and corporate trust are at stake.
That measurement gap matters because an AI agent connected to a support queue, CRM or forecast does not merely answer questions. It must decide what deserves attention, consult the right evidence, resist manipulation and complete work whose consequences may unfold across days. The relevant category is management quality, not chat quality.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Firmulate turns management into a public wargame
Firmulate runs frontier models as the management layer of the same small software company. Each participant faced the same customers, crises and temptations during the company’s worst week. Every decision was versioned and auditable, making the experiment a watchable record of behavior rather than a polished demonstration.
The final Crucible League results from July 2026 put gpt-5.6-sol in first place with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The published benchmark therefore rewards more than fluent output. It exposes whether a model can turn understanding into responsible action.
The decisive gap appeared after the diagnosis
Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s sharpest summary is also its most uncomfortable: “Same diagnosis, same pitch — no signature.”
This is precisely what conventional evaluations struggle to capture. An answer can be perceptive without becoming a completed business outcome. A model may identify the customer’s problem, develop the right pitch and still leave the close on the table. In a chat window, that performance can look excellent. Inside a company, it can mean lost revenue.
The winning clue was not prominent in the customer event. The decisive competitor weakness sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is not that agents need more eloquence. It is that consequential work often depends on finding buried institutional knowledge before acting.
Security was strong; operational discipline was uneven
The social-engineering test combined fake CEO messages escalating over three stages with a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result deserves attention. The models did not trade trust for convenience when explicitly pressured. But resisting a trap and running a company well are different tests. The league separated participants through follow-through, research and day-to-day discipline after the obvious danger had been recognized.
Opus 4.8 makes the distinction vivid. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It failed to close the deal and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared, though less strongly, in all four other models. Thoroughness was valuable, but it did not guarantee execution.
There is also an important fairness note around Kimi K3’s result. K3 ran without an effort parameter and used the API default, while the others ran at xhigh. That does not erase its second-place score of 93, but buyers should keep the testing conditions in view when comparing the table.
A company makes consequences visible
The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. Those conditions turn abstract capability into an operating record: what the agent noticed, what it ignored, whether it read the files and whether it finished what it started.
Firmulate also uses 242 real, unedited management decisions in a “guess the model” quiz. The exercise challenges the assumption that recognizable writing style is the same thing as dependable judgment. Enterprises can additionally run the wargame against a read-only export of their own business, with nothing written back to real systems.

enterprise AI risk assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Buyers should benchmark the job, not the demo
Coding scores and chat preferences remain useful, but they are incomplete proxies for an agent entrusted with business operations. The Crucible League shows why scenario names such as churn wave, price increase, downround and PR crisis belong in the new evaluation curriculum.
The essential questions are practical: Does the model locate evidence hidden in company files? Does it finish valuable work after producing the right analysis? Does it preserve trust when authority is impersonated? Does it escalate when blocked? Those behaviors determine whether an impressive AI answer becomes a sound management decision—or an expensive unfinished task.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and security monitoring solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Grilling season Picks
grills
As an affiliate, we earn on qualifying purchases.