TL;DR
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
Create a free accountAs an affiliate, we earn on qualifying purchases.
A live experiment by Firmulate tested AI models in managing a small company’s worst week, revealing that management skills, not just chat quality, are crucial. The results show models can identify crises but often fail at execution and trust, emphasizing a new evaluation approach.
AI models’ ability to manage complex, real-world business crises was tested in a live experiment by Firmulate, revealing critical gaps beyond traditional chat and coding benchmarks. The experiment involved five AI managers overseeing a small company’s worst week, with results showing that while models can diagnose crises, they often fail to execute decisions effectively or maintain trust. This development underscores a need to evaluate AI performance in practical management tasks, not just conversational or technical accuracy. For more on this, see the original analysis here.
The Firmulate experiment placed five AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—in a simulated business environment where they managed a company facing multiple crises, including customer churn, PR issues, and financial decisions. Details can be found in the original analysis. The models were scored on their ability to diagnose problems, communicate effectively, and make decisions under pressure. The top performer, gpt-5.6-sol, achieved a score of 95, while Opus 4.8 scored 73. Despite all models identifying crises and resisting manipulation attempts, only two signed a significant deal, illustrating that diagnosis alone is insufficient for successful management.One key finding was that models could sound informed but often failed to retrieve or act on the critical piece of information that would influence business outcomes. For example, a model that read the company’s files but did not present a crucial reference lost a €55,000 deal, highlighting a gap between understanding and execution. Furthermore, models refused social engineering attempts, demonstrating safety features, but still struggled with completing managerial tasks such as escalation and follow-through. The experiment’s design enforced strict trust standards, with breaches capping the overall score, emphasizing the importance of honesty and reliability in management AI.
Implications for AI Management Evaluation
The results from the Firmulate experiment suggest that current AI benchmarks, focused on chat or coding, do not adequately measure management capabilities. The ability to diagnose crises, communicate effectively, and execute decisions reliably is vital for deploying AI in real-world business contexts. The experiment highlights that AI’s management skills are a distinct category requiring dedicated evaluation, especially as organizations consider AI for operational decision-making, customer relations, and strategic planning. The findings stress that safety, trust, and execution are equally critical as informational accuracy, shaping future standards for AI assessment in enterprise settings.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limitations of Traditional AI Benchmarks in Business Management
Traditional AI evaluation metrics often focus on technical performance, such as coding accuracy or conversational quality, which do not reflect real-world management challenges. The Firmulate experiment builds on prior work showing that models excel at isolated tasks but falter in integrated, consequence-driven environments. The July 2026 Crucible League was designed to simulate a company’s worst week, incorporating real money mechanics, versioned work, and decision tracking, creating a rigorous test bed for management skills. This approach aims to bridge the gap between benchmark scores and actual business performance, emphasizing the importance of trust, escalation, and long-term decision quality.
“The real test for AI managers is whether they can manage consequences, not just produce correct answers.”
— Thorsten Meyer, founder of Firmulate
business crisis management AI software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Challenges in AI Management Testing
While the Firmulate experiment provides valuable insights, it remains unclear how well these results generalize to different industries or larger organizations. The simulation was designed with a specific small business scenario, and real-world complexity varies widely. Additionally, the long-term impact of integrating AI managers into operational workflows, including issues of trust, accountability, and human-AI collaboration, is still being studied. Future research is needed to refine evaluation metrics, explore broader scenarios, and understand how AI can reliably handle unpredictable, high-stakes environments.
AI management simulation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarks
Following the July 2026 results, Firmulate plans to expand its testing framework to include larger organizations and more diverse scenarios, such as crisis escalation, multi-team coordination, and strategic decision-making. Industry stakeholders are encouraged to adopt similar live, consequence-based testing environments to evaluate AI tools beyond traditional benchmarks. The goal is to develop standardized metrics for management performance, emphasizing trustworthiness, execution, and long-term impact. As AI models improve, ongoing assessments will be critical to ensure they can manage real-world complexity without compromising safety or operational integrity.
enterprise AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What makes the Firmulate experiment different from traditional AI benchmarks?
The experiment tests AI models in managing a simulated business crisis, focusing on decision-making, execution, trust, and consequences, rather than just chat quality or coding accuracy.
Can AI models reliably manage real business operations based on these results?
The results show promise but also reveal significant gaps in execution and trust. More testing and development are needed before deploying AI as operational managers.
What are the main weaknesses identified in the models?
Models often fail to follow through on decisions, retrieve critical information, or escalate issues properly, despite diagnosing problems accurately.
Will this lead to new standards for AI evaluation?
Yes, the findings suggest a need for management-specific benchmarks that measure decision quality, trustworthiness, and long-term outcomes in real-world scenarios.
How can organizations prepare for AI management integration?
Organizations should consider live, consequence-based testing environments to assess AI tools’ management capabilities and ensure alignment with operational trust and safety standards.
Source: ThorstenMeyerAI.com
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.