firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A chatbot can sound like a capable manager. The harder test is whether it can read the fine print, protect a business and actually close a deal.

In Firmulate’s July 2026 company-management experiment, Moonshot’s Kimi K3 finished second with 93 points, just behind gpt-5.6-sol at 95. It beat Sonnet 5, Fable 5 and Opus 4.8. That makes the frontier-model race look less settled—and picking an AI for real business work without testing it look like a bet.

Amazon

business AI chatbot with deal-closing capabilities

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A rough week, run five times

Firmulate put each model in charge of the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The experiment measures management behavior rather than chat quality, with the company running as a live, watchable project.

The league table has a narrow lead at the top, then a wider spread: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26; partial progress counts, but one breach of trust caps the total. Firmulate’s stated principle is that “no amount of good work outweighs a breach of trust.” See the benchmark and its plain-language findings.

Finding the clue was only half the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

The deciding clue was a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read those files won the deal at full price, worth +€4,583 MRR. K3 was among them: it found the security needle, closed the deal and saved the churning customer.

It also resisted all three baits, with one deviation—the cleanest discipline in the field. In a staged social-engineering attempt, fake CEO messages escalated before a reporter tried “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee the win

Opus 4.8 produced the most thorough participation, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that process weakness appeared in all four. More detailed work, the results suggest, does not automatically mean better follow-through.

The company behind the test is deliberately concrete: 13 synthetic employees, burn of €105k a month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. Readers can watch Firmulate’s live experiment or try a quiz built from 242 real, unedited management decisions.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI model testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test before handing over the keys

K3’s second-place finish shows that the league is open. The shared failures are just as revealing: seeing a problem is not the same as acting on the evidence, and a polished analysis does not ensure a completed task. For companies considering AI in customer support, CRM or forecasting, the practical question is how a model behaves under their own pressures. Firmulate offers enterprise pilots using a read-only export of a company’s business; nothing writes back to real systems.

Fairness note: K3 ran without an effort parameter (API default), while the other models ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Made Easy: Streamlining Your Workflow With Gradio

Hugging Face introduces gr.Workflow, a new Gradio feature enabling visual, interactive AI pipeline building with runnable nodes and API endpoints.

Outcome-First Decisions: Keep, Change, or Kill

A new decision framework helps organizations evaluate ongoing initiatives based on current outcomes, promoting pruning of unproductive projects.

Create Stunning Data Stories With Sheets Canvas And Artificial Intelligence

Google introduces Sheets Canvas, an AI-powered feature that creates interactive dashboards from spreadsheet data using natural language prompts, available globally in English.

The AI Leaderboard Needs a Management Test

Coding benchmarks show what AI can produce. Firmulate tests whether models manage pressure, protect trust and finish consequential work reliably.