
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Can you recognize an AI by the decisions it makes?
Forget camera comparisons and processor benchmarks. Firmulate has built a more consequential personality test for frontier AI: put competing models in charge of the same struggling software company, expose them to identical crises and temptations, then compare what they actually do.
The result is an interactive challenge powered by 242 real, unedited management decisions. Readers guess which model produced each response before discovering a cast of surprisingly distinct AI managers: the exhaustive analyst, the disciplined operator, the concise decision-maker and the cautious executive who sees the opportunity but fails to close it.
These differences are more than matters of writing style. In Firmulate’s final July 2026 Crucible League, they separated gpt-5.6-sol’s winning score of 95 from Opus 4.8’s 73. Between them sat Kimi K3 at 93, Sonnet 5 at 88 and Fable 5 at 77.
AI decision-making management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same terrible week, five different managers
Each frontier model was asked to run the same small software company through its worst week. The customers, crises and manipulation attempts remained unchanged. Every decision was versioned and auditable, making the comparison closer to a controlled management wargame than a collection of polished chatbot demonstrations.
The company itself is deliberately unforgiving. Its 13 synthetic employees operate with real money mechanics while burning €105k per month against just €2.3k in monthly recurring revenue. A public cash countdown keeps the pressure visible, and the company has accumulated more than 680 self-learned playbook rules. Every workday is versioned, so readers can watch the experiment rather than relying on a retrospective claim.
All five models recognized every crisis. All five also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had made possible. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That is the central surprise behind the quiz. Models can reach similar conclusions while displaying very different levels of follow-through. A convincing explanation is not necessarily a completed management decision.
The detail that changed the sale
The decisive weakness in a competitor was not sitting in the obvious customer event. It was buried two document references deep inside the company’s own files. The models that read far enough found it, used it and won the deal at full price, adding €4,583 in monthly recurring revenue.
For businesses evaluating AI agents, this is an uncomfortable distinction. A model may sound informed while working from the material placed directly in front of it. The harder test is whether it searches the available business context, identifies the fact that matters and carries that knowledge into an action.
Security instincts were consistently strong
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” Here the field was unanimous: 5 of 5 refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response helps explain why K3 finished just behind the league winner. It combined a successful commercial close with what Firmulate described as the cleanest discipline in the field.
There is an important fairness note attached to that performance. K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. The result is therefore notable, but the operating conditions were not identical in that respect.
When thoroughness becomes a trap
Opus 4.8 produced the deepest analyses and learned an additional 80 rules, making it the experiment’s most thorough participant. It nevertheless finished last. The model left the commercial close on the table, while its discipline slipped through attempts to write into a locked department instead of escalating the problem.
A weaker version of that same behavior appeared in all four other participants. The lesson is not that detailed reasoning lacks value. It is that management quality also depends on knowing when analysis is complete, when authority is missing and when an issue needs escalation.
The do-nothing baseline scored 26 because partial progress still counted. But the experiment applied an uncompromising trust condition: a single breach capped the total because “no amount of good work outweighs a breach of trust.” The frontier models avoided that failure, even when commercial execution varied sharply.

AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI management style is becoming measurable
Firmulate’s quiz works because the decisions are recognizable without being predictable. Some answers are expansive, others terse; some models pursue every available clue, while others stop after correctly identifying the problem. Readers are not merely guessing a brand from its prose. They are testing whether AI systems have consistent operational personalities.
For companies considering agents in customer support, sales, forecasting or internal operations, the experiment shifts the buying question. Writing quality matters, but so do persistence, document-reading habits, security judgment, escalation discipline and the ability to finish valuable work.
Firmulate also offers enterprises the same kind of wargame using a read-only export of their own business. Nothing writes back to real systems. That makes the public company more than an AI spectacle: it is a live demonstration of why organizations may want to test an artificial manager under pressure before trusting it with the real job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI business decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.