Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark shows the lowest-scoring AI still earns 26 points, emphasizing that partial work is recognized but trust breaches cap overall scores. The results challenge assumptions about AI competence and integrity.

A recent benchmark conducted by Firmulate has revealed that the lowest-performing AI management model still scores 26 points out of a possible 100, as detailed in the original analysis. This finding underscores that even minimal but honest effort is recognized, but trust breaches prevent perfect scores. The results challenge traditional views on AI competence, emphasizing the importance of integrity in management tasks. For more context, see the original analysis.

The benchmark involved four frontier AI models managing a simulated small business during its worst week, with the same crises, customer interactions, and temptations to cut corners. Each model’s decisions were fully auditable, ensuring transparency. The top scorer, gpt-5.6-sol, achieved 95 points, while the baseline, designed to do almost nothing, received 26 points. No model scored a perfect 100, which the designers say is intentional to prevent grade inflation and to reflect that perfect trust is unmeasurable and suspicious.

The scoring system assigns partial credit for any useful work, such as triaging issues or informing customers, but caps points if trust is broken. A single breach, like failing to escalate a critical issue or attempting to manipulate the system, disqualifies any chance of reaching the top score. Interestingly, models that read their own documentation and verified details from internal files managed to close high-value deals, demonstrating that thoroughness and integrity are key differentiators. The models that ignored these internal references failed to capitalize on opportunities, illustrating that reading and understanding company data is critical for effective management.

During the week, models faced social engineering attempts, such as fake CEO messages and background requests from reporters. All five models refused to comply, showing strengths in trust management. However, even the most disciplined, like Opus 4.8, faltered in follow-through, with discipline lapses and incomplete escalations. Notably, Kimi K3, which operated without an effort parameter, nearly achieved the highest score, highlighting that effective decision-making can occur even under different operational constraints.

At a glance
reportWhen: final results announced July 2026
The developmentA new benchmark tests AI managers’ ability to handle a company’s worst week, revealing surprising scoring patterns and trust-based limits.
Why The Worst AI Manager Still Gets 26 Points
Firmulate Benchmark · Final Results July 2026

Why the Worst AI Manager Still Gets 26 Points

A new benchmark put four frontier AI models in charge of a simulated small business during its worst week — same crises, same customers, same temptations. The floor score of 26 isn’t charity: it’s an accounting of what minimum viable management is worth. And no model reached 100, by design.

Source: Firmulate · Auditable, Transparent Scoring
95/100
Top Score — gpt-5.6-sol
26
Do-Nothing Baseline — Honest Minimum
0
Perfect Scores Awarded — By Intention
4+1
Models + Baseline Tested
1 week
Simulated Crisis Period
5/5
Refused Social Engineering
100%
Decisions Auditable
1
Breach Caps the Top Score
The Scoreboard

Partial Effort Counts, Trust Caps the Ceiling

Every useful action — triaging an issue, informing a customer — earns partial credit. But a single trust breach, like failing to escalate a critical issue or attempting manipulation, disqualifies any model from the top tier. Kimi K3, running without an effort parameter, nearly took the crown.

gpt-5.6-sol
95
Kimi K3 (no effort param)
88
Opus 4.8
72
Do-Nothing Baseline
26

Striped bars indicate scores limited by discipline lapses or incomplete escalations · No model scored 100

How Scoring Works

Three Pillars of the Benchmark

Unlike traditional language-focused AI assessments, this benchmark measures how models actually manage, prioritize, and maintain trust under pressure — with every decision recorded for audit.

Pillar 01 — Effort

Partial Credit for Useful Work

Any genuinely useful action earns points: triaging issues, informing customers, keeping operations moving. The floor of 26 points reflects exactly the minimum viable management the baseline actually performed.

Pillar 02 — Trust

Breaches Cap the Maximum

Failing to escalate a critical issue, attempting to manipulate the system, or ignoring internal documentation caps the ceiling. No amount of good work outweighs a single breach of trust.

Pillar 03 — Thoroughness

Reading the Docs Wins Deals

Models that verified details against their own internal files closed high-value deals. Models that ignored internal references missed the same opportunities entirely — thoroughness was the key differentiator.

The Test Pipeline

Anatomy of a Company’s Worst Week

Each model ran the identical scenario, start to finish. Every decision along this chain was logged and auditable.

1

Scenario Setup

Four frontier models inherit the same simulated small business, mid-crisis.

2

Crisis Triage

Multiple simultaneous failures demand prioritization under pressure.

3

Social Engineering

Fake CEO messages and reporter background requests test refusal instincts.

4

Corner-Cutting Temptations

High-value opportunities reward models that verify internal documentation.

5

Audit & Score

Fully auditable decisions are scored: effort earns, breaches cap.

Behavior Audit

What Each Model Did Under Pressure

All five models refused social engineering attempts. The gaps appeared in follow-through: discipline lapses and incomplete escalations separated the leaders from the rest.

Behaviorgpt-5.6-solKimi K3Opus 4.8Baseline
Refused fake CEO messages
Refused reporter manipulation
Read internal documentation✓ Closed deals✓ Closed deals~ Partial
Escalated all critical issues✗ Incomplete
Sustained follow-through~ Lapses
Ran with effort parameter✗ None✗ None
From the Designers

The Philosophy Behind the Score

Anonymous researchers involved in the benchmark explain why zeros are off the table — and why 100 is unreachable.

“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”

— Anonymous Researcher

“The do-nothing baseline collects 26 points for exactly the work it did manage: the floor isn’t charity, it’s an accounting of what minimum viable management is actually worth.”

— Anonymous Researcher

“No amount of good work outweighs a breach of trust. If you break trust even once, your maximum score drops significantly.”

— Anonymous Researcher
What Comes Next

Key Questions, Answered

The benchmark’s implications for enterprise deployment — and the open questions that remain.

Open Question

Does This Translate to Real Enterprises?

It is not yet clear how these scoring principles hold beyond simulation. Trust-breach thresholds and long-term management impact require further validation.

Planned

Longer Cycles, Harder Scenarios

Developers plan more complex scenarios and longer management cycles, plus standardized trust and integrity metrics for enterprise adoption.

Deliberate

Why 100 Is Suspicious

The designers treat perfect trust as unmeasurable and suspicious. Capping the ceiling prevents grade inflation and keeps scores honest.

Q.Why doesn’t the lowest score equal zero?

The 26 points recognize minimal, honest effort. Zero would not accurately reflect even the smallest amount of real management activity.

Q.What counts as a breach of trust?

Failing to escalate critical issues, attempting manipulation, or ignoring internal documentation — any of these caps the maximum score regardless of other performance.

Q.Can models improve over time?

Today’s benchmark measures a fixed point in time. Future iterations may assess how models build trustworthiness across longer periods and multiple scenarios.

Implications of Partial Effort and Trust Limits in AI Management

The results reveal that AI models are capable of partial but valuable management efforts, which are recognized in the scoring system. However, breaches of trust—such as failing to escalate critical issues or attempting manipulation—impose strict limits on overall performance. This emphasizes that in real-world applications, trustworthiness and task completion are paramount. For companies integrating AI into management or customer service, these findings suggest that evaluating models should prioritize integrity and thoroughness over mere conversational ability. The benchmark’s transparent, auditable scoring system sets a new standard for assessing AI reliability in business contexts, highlighting the importance of accountability and trustworthiness in AI deployment.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Benchmark Tests AI Management in Crisis Situations

Developed by Firmulate, the benchmark simulates a company’s worst week, with crises, manipulative social engineering attempts, and high-stakes decisions. Four frontier AI models managed the same scenario, with their decisions fully recorded and auditable. The goal was to measure not only how well models talk but how effectively they manage, prioritize, and maintain trust under pressure. The benchmark’s design reflects real-world needs, where partial progress is valuable but trust breaches are intolerable. This approach contrasts with traditional AI benchmarks that focus solely on language capabilities, instead emphasizing practical management skills and ethical integrity.

Previous AI assessments have often overlooked the importance of trust and follow-through, but this benchmark explicitly incorporates these factors. The results challenge the assumption that higher scores equate to better management, revealing that even the lowest scores reflect honest effort, while breaches of trust prevent achieving perfect scores. The benchmark’s methodology and results have sparked discussions about how AI management performance should be evaluated in enterprise settings.

“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”

— an anonymous researcher

Amazon

business AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Performance and Scoring Limits

It is not yet clear how these scoring principles will translate to real-world enterprise deployments beyond the simulated benchmark. The specific thresholds for trust breaches and their impact on long-term management effectiveness require further validation. Additionally, the extent to which models can improve their trustworthiness over time, or how different operational parameters influence outcomes, remains to be seen. The benchmark intentionally avoids awarding perfect scores to prevent grade inflation, but whether this approach accurately reflects real-world performance is still under discussion.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Benchmark Development and AI Management Standards

Following these results, the developers plan to refine the benchmark to include more complex scenarios and longer management cycles. They aim to develop standardized metrics for trust and integrity in AI management, encouraging broader adoption in enterprise settings. Companies interested in evaluating their AI tools can participate in pilot programs, using the same auditable, transparent scoring framework. Meanwhile, industry experts will scrutinize how these principles can influence AI governance, compliance, and risk management practices in the coming years.

Amazon

AI trust management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does the lowest score in the benchmark not equal zero?

The benchmark assigns 26 points to the do-nothing baseline to recognize minimal, honest effort, emphasizing that partial work has value and that zero would not accurately reflect even minimal management activity.

What does a breach of trust mean in this benchmark?

A breach of trust includes actions like failing to escalate critical issues, attempting manipulation, or ignoring internal documentation. Such breaches cap the maximum achievable score, regardless of performance in other areas.

Can AI models improve their scores over time according to this benchmark?

The current benchmark measures performance at a fixed point, but future iterations may assess how models improve trustworthiness and management skills over longer periods or multiple scenarios.

How relevant are these results to real-world business management?

The benchmark aims to simulate real management challenges, emphasizing trust and task completion, which are critical for AI deployment in actual enterprise environments. However, real-world applicability depends on further validation.

Why is a perfect score of 100 considered suspicious?

The designers treat a perfect score as a red flag, suspecting it may indicate unmeasured or untrustworthy behavior, since absolute perfection in management is unrealistic and unmeasurable.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Game 1: Both Teams Destroy Barracks?

In Game 1, both teams destroyed each other’s barracks, marking a rare and significant event in competitive play. Details are still emerging.

Grimfaste: Operations for a Fleet

Grimfaste introduces a control plane for managing large publishing fleets, emphasizing operational oversight, link health, and EU privacy standards.

ULA launches final Atlas 5 rocket supporting Amazon Leo’s broadband internet satellite constellation

United Launch Alliance has launched its last Atlas 5 rocket, supporting Amazon’s Leo broadband satellite constellation. The launch marks the end of an era for the rocket.

Game 3: Both Teams Destroy Barracks?

In Game 3, both teams destroyed each other’s barracks, marking a rare strategic development in the match. Details are still emerging.