
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A chatbot can sound like a capable manager. The harder test is whether it can read the fine print, protect a business and actually close a deal.
In Firmulate’s July 2026 company-management experiment, Moonshot’s Kimi K3 finished second with 93 points, just behind gpt-5.6-sol at 95. It beat Sonnet 5, Fable 5 and Opus 4.8. That makes the frontier-model race look less settled—and picking an AI for real business work without testing it look like a bet.
business AI chatbot with deal-closing capabilities
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A rough week, run five times
Firmulate put each model in charge of the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The experiment measures management behavior rather than chat quality, with the company running as a live, watchable project.
The league table has a narrow lead at the top, then a wider spread: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26; partial progress counts, but one breach of trust caps the total. Firmulate’s stated principle is that “no amount of good work outweighs a breach of trust.” See the benchmark and its plain-language findings.
Finding the clue was only half the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
The deciding clue was a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read those files won the deal at full price, worth +€4,583 MRR. K3 was among them: it found the security needle, closed the deal and saved the churning customer.
It also resisted all three baits, with one deviation—the cleanest discipline in the field. In a staged social-engineering attempt, fake CEO messages escalated before a reporter tried “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness did not guarantee the win
Opus 4.8 produced the most thorough participation, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that process weakness appeared in all four. More detailed work, the results suggest, does not automatically mean better follow-through.
The company behind the test is deliberately concrete: 13 synthetic employees, burn of €105k a month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. Readers can watch Firmulate’s live experiment or try a quiz built from 242 real, unedited management decisions.

As an affiliate, we earn on qualifying purchases.
Test before handing over the keys
K3’s second-place finish shows that the league is open. The shared failures are just as revealing: seeing a problem is not the same as acting on the evidence, and a polished analysis does not ensure a completed task. For companies considering AI in customer support, CRM or forecasting, the practical question is how a model behaves under their own pressures. Firmulate offers enterprise pilots using a read-only export of a company’s business; nothing writes back to real systems.
Fairness note: K3 ran without an effort parameter (API default), while the other models ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
