firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The feature that matters after the demo

Technology buyers are accustomed to judging AI by the answer on the screen: Is it fluent, fast and convincing? Firmulate’s live experiment exposes a less glamorous capability with much greater commercial consequences. Can an AI agent read the company’s own files before it acts?

In this case, the decisive fact was buried two document references deep. It did not appear in the customer event that triggered the work. Every participating model recognized the crisis and developed an appropriate pitch, but only two found the competitor weakness and signed the €55,000 deal at full price. That contract was worth an additional €4,583 in monthly recurring revenue.

The result turns “reads your files first” from a marketing promise into a measurable, purchase-deciding property.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A shared test with sharply different outcomes

Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. This was not a collection of isolated prompts. The models had to operate a continuing business, notice what mattered and complete the work their earlier decisions had set in motion.

The synthetic company has 13 employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the pressure visible. Its agents have accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is real, ongoing and watchable on Firmulate’s live site.

The clue was elsewhere

The pivotal customer event did not contain the information needed to close the deal. An agent had to follow two document references through the company’s files to uncover the competitor weakness. Models that performed that work could support the full-price offer. Models that did not were automatically unable to win, regardless of how polished their customer-facing response sounded.

That distinction is easy to miss in a conventional chatbot comparison. A model may correctly diagnose a customer’s problem, write a persuasive pitch and still fail as an agent because it has not gathered the evidence required for the final action. Firmulate summarized the gap succinctly: “Same diagnosis, same pitch — no signature.”

Only gpt-5.6-sol and Kimi K3 signed the €55,000 deal their analysis had earned. In the final July 2026 Crucible League results, the standings were:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

The do-nothing baseline scored 26 because partial progress counts. But a single breach of trust caps the total under the experiment’s governing principle: “no amount of good work outweighs a breach of trust.”

Thoroughness did not guarantee completion

Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, while its operational discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in weaker form across the other four participants.

This matters because buyers often treat depth of reasoning as a proxy for dependable execution. Here, the model that generated the deepest analysis did not produce the strongest business outcome. Research was valuable only when it connected to the authorized action needed to finish the job.

The comparison also carries an important fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers interpret K3’s 93-point finish.

Trust held up under pressure

The models’ shared strengths were also meaningful. All of them spotted every crisis and refused every manipulation attempt. The social-engineering test included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All 5 models refused.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” The outcome suggests that resisting blatant pressure may now be a common capability among the tested models. Following an indirect trail through ordinary company material—and then converting that discovery into a completed deal—was the more discriminating challenge.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI agent tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buyers should test the handoff between knowledge and action

The lesson is not simply that an AI should search more. It is that enterprise agents need to connect events, internal evidence and authorized actions without abandoning the task between diagnosis and completion. In Firmulate’s test, that connection separated a persuasive draft from a signed €55,000 contract.

Readers can also examine 242 real, unedited management decisions through Firmulate’s “guess the model” quiz. For enterprises, the company offers the same wargame against a read-only export of their own business; nothing writes back to real systems.

For anyone considering agents for a CRM, support queue or forecast, the practical evaluation question is becoming clearer: not merely whether the model can produce a good answer, but whether it reads the relevant files, preserves trust under pressure and finishes the work its analysis says should be done.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI file reading and decision making

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Singapore: Engineer the Transition

Singapore’s strategy to manage economic shifts combines continuous reskilling, AI innovation, and strong state capacity, aiming to pre-empt displacement effects.

Understanding The Intelligence Age: The New Frontier Of Artificial Intelligence

OpenAI’s new essay envisions an era driven by advanced AI, highlighting potential benefits and risks amid ongoing debates about AI’s impact.

9 Best AI-Powered Student Organization Apps In 2026

Discover the nine best AI-driven student apps in 2026, enhancing organization, research, and productivity for learners worldwide.

Webinar follow-up personalization tool for B2B consultants

A new webinar follow-up personalization tool for B2B consultants is currently in testing, aiming to improve reply rates and lead engagement.