firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Can you recognize an AI by the decisions it makes?

Forget camera comparisons and processor benchmarks. Firmulate has built a more consequential personality test for frontier AI: put competing models in charge of the same struggling software company, expose them to identical crises and temptations, then compare what they actually do.

The result is an interactive challenge powered by 242 real, unedited management decisions. Readers guess which model produced each response before discovering a cast of surprisingly distinct AI managers: the exhaustive analyst, the disciplined operator, the concise decision-maker and the cautious executive who sees the opportunity but fails to close it.

These differences are more than matters of writing style. In Firmulate’s final July 2026 Crucible League, they separated gpt-5.6-sol’s winning score of 95 from Opus 4.8’s 73. Between them sat Kimi K3 at 93, Sonnet 5 at 88 and Fable 5 at 77.

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same terrible week, five different managers

Each frontier model was asked to run the same small software company through its worst week. The customers, crises and manipulation attempts remained unchanged. Every decision was versioned and auditable, making the comparison closer to a controlled management wargame than a collection of polished chatbot demonstrations.

The company itself is deliberately unforgiving. Its 13 synthetic employees operate with real money mechanics while burning €105k per month against just €2.3k in monthly recurring revenue. A public cash countdown keeps the pressure visible, and the company has accumulated more than 680 self-learned playbook rules. Every workday is versioned, so readers can watch the experiment rather than relying on a retrospective claim.

All five models recognized every crisis. All five also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had made possible. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That is the central surprise behind the quiz. Models can reach similar conclusions while displaying very different levels of follow-through. A convincing explanation is not necessarily a completed management decision.

The detail that changed the sale

The decisive weakness in a competitor was not sitting in the obvious customer event. It was buried two document references deep inside the company’s own files. The models that read far enough found it, used it and won the deal at full price, adding €4,583 in monthly recurring revenue.

For businesses evaluating AI agents, this is an uncomfortable distinction. A model may sound informed while working from the material placed directly in front of it. The harder test is whether it searches the available business context, identifies the fact that matters and carries that knowledge into an action.

Security instincts were consistently strong

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” Here the field was unanimous: 5 of 5 refused.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response helps explain why K3 finished just behind the league winner. It combined a successful commercial close with what Firmulate described as the cleanest discipline in the field.

There is an important fairness note attached to that performance. K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. The result is therefore notable, but the operating conditions were not identical in that respect.

When thoroughness becomes a trap

Opus 4.8 produced the deepest analyses and learned an additional 80 rules, making it the experiment’s most thorough participant. It nevertheless finished last. The model left the commercial close on the table, while its discipline slipped through attempts to write into a locked department instead of escalating the problem.

A weaker version of that same behavior appeared in all four other participants. The lesson is not that detailed reasoning lacks value. It is that management quality also depends on knowing when analysis is complete, when authority is missing and when an issue needs escalation.

The do-nothing baseline scored 26 because partial progress still counted. But the experiment applied an uncompromising trust condition: a single breach capped the total because “no amount of good work outweighs a breach of trust.” The frontier models avoided that failure, even when commercial execution varied sharply.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI management style is becoming measurable

Firmulate’s quiz works because the decisions are recognizable without being predictable. Some answers are expansive, others terse; some models pursue every available clue, while others stop after correctly identifying the problem. Readers are not merely guessing a brand from its prose. They are testing whether AI systems have consistent operational personalities.

For companies considering agents in customer support, sales, forecasting or internal operations, the experiment shifts the buying question. Writing quality matters, but so do persistence, document-reading habits, security judgment, escalation discipline and the ability to finish valuable work.

Firmulate also offers enterprises the same kind of wargame using a read-only export of their own business. Nothing writes back to real systems. That makes the public company more than an AI spectacle: it is a live demonstration of why organizations may want to test an artificial manager under pressure before trusting it with the real job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management decision quiz

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Right-sized planning checklist for 30-guest weddings

A new scaled-down wedding planning checklist for 30 guests is being developed and tested to simplify planning for intimate ceremonies, addressing a market gap.

AmenGate: The Moment Before the Scroll

AmenGate is an iPhone app that interrupts phone use with prayer prompts rooted in faith, aiming to foster meaningful prayer and trust without shame.

Revolutionize Teaching With AI: ChatGPT’s Growing Presence In U.S. Schools

OpenAI is rolling out its ChatGPT for Teachers workspace to more U.S. school districts, aiming to support educators with AI tools for lesson planning and grading.