When An AI Agent’s Claim Of Completion Meets The Database
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: When An AI Agent’s Claim Of Completion Meets The Database on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Microsoft and Hugging Face have made ThinkingBox available through Hugging Face. The benchmark checks whether AI agents leave business systems in required database states across 507 workflows, with 20 runs per task; its authors report substantial failures in their tested setup.

Microsoft and Hugging Face have made ThinkingBox, a benchmark for AI agents, available through Hugging Face. It tests whether agents leave business systems in the required database state—not simply whether they make valid tool calls or provide plausible answers—across 507 workflows run 20 times each, extending the original analysis of agents and database state.

ThinkingBox runs agents in isolated sessions using tools built on the Model Context Protocol (MCP), then checks the resulting backend records and side effects against executable requirements. The release says the workflows cover retail, auto insurance, travel, neobanking and consulting. The benchmark can be run through OpenEnv.

In a common-set analysis, the authors report 121,680 valid trials across 12 models, of which 79,853 failed executable checks. Among those failed attempts, 67.24% ended without a final tool error despite the agent having invoked a state-changing tool. The checks identified incorrect field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%; those categories overlap.

The release uses several measures: pass@1 is the share of attempts that succeed, pass@20 indicates whether a task succeeded at least once in 20 runs, and observed 20/20 counts tasks that passed all 20 recorded runs. The authors report pass@1 scores of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, which they describe as the strongest open-weight model in their table. They say Kimi-K3 scored within one percentage point of GPT-6 Astra.

At a glance
announcementWhen: Availability announced; the supplied so…
The developmentMicrosoft and Hugging Face have released ThinkingBox, a benchmark that evaluates AI agents by checking the backend state and side effects their workflow actions produce.
At a glance
announcementWhen: Now available through Hugging Face; the…
The developmentMicrosoft and Hugging Face released ThinkingBox through Hugging Face, a benchmark for evaluating AI agents by their backend changes across repeated workflow trials.

Why Database Outcomes Matter

The benchmark focuses on a gap between an agent appearing to complete a task and the requested change actually being recorded. In business workflows, the difference can affect whether a support case is handled correctly, a payment is issued, or a booking or claim is updated as required. A fluent response does not by itself show that the underlying system has the right state.

Repeated runs address another deployment concern: consistency. An agent that succeeds once may fail on another attempt at the same task. Pass@1, pass@20 and observed 20/20 describe different aspects of performance, but the authors’ results are measurements within their benchmark setup, not guarantees about an organization’s live systems. Teams can use the checks to find workflow failures and compare models, while still needing tests that reflect their own records, policies and integrations.

Amazon

database testing tools for AI workflows

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Tool Calls to System State

Many evaluations can focus on whether an agent selected an available tool, completed a sequence of calls or produced a suitable final response. ThinkingBox instead checks the state left behind and any additional effects. Its authors’ stated premise is that valid tool use does not prove the task’s requirements were met.

The release illustrates this with a delayed $745 appliance order. An agent investigates the delay, opens a support ticket and records a timeline. The customer is not eligible for late-delivery compensation under the policy the agent checked, but the carrier exception remains open and the task requires the ticket to stay on hold. The agent marks it solved and replies without answering the customer’s underlying question. The executable check fails because the ticket is solved rather than on hold.

The supplied source says the release is based on the authors’ paper, but does not provide a publication date. It also does not include the full model table or all details of the evaluation setup.

“A tool call is not an outcome.”

— Microsoft and Hugging Face, in the ThinkingBox release

Amazon

business workflow automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Benchmark Results

The reported scores and failure categories are the authors’ results in the tested setup. The supplied material does not provide uncertainty estimates for the model scores, full task specifications, or all model configurations. It therefore does not support treating small score differences as definitive rankings.

It is also unclear how well performance transfers to live business systems with changing records, unusual requests, different policies or integrations not represented in the benchmark. Passing all 20 recorded runs is evidence limited to those observations; the source does not establish that it predicts long-term reliability. No independent replication or future release milestone is identified in the supplied material.

Amazon

AI system monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Running Tests on Real Workflows

The release says developers and researchers can run ThinkingBox through OpenEnv with isolated MCP tool sessions. That enables examination of the tasks and comparison of agent outcomes against executable state checks. The source does not announce a scheduled next milestone or provide a date for further updates.

For organizations considering agents, the practical next step is to test workflows against their own requirements and inspect both successful and failed runs. ThinkingBox offers a way to measure recorded outcomes and repeatability in its test environment; whether those measurements predict performance in a particular organization remains an open question.

Amazon

backend database validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does ThinkingBox evaluate?

It checks whether an AI agent leaves backend records and side effects in the required state after completing a workflow.

How large is the benchmark?

The release describes 507 business workflows, with each task repeated 20 times. The workflows span retail, auto insurance, travel, neobanking and consulting.

What did the authors report about failures?

In a common-set analysis of 121,680 valid trials across 12 models, the authors say 79,853 attempts failed executable checks. They report overlapping failure categories, including incorrect field values, unintended extra effects and missing required effects.

Do the scores predict how agents will perform at a company?

Not by themselves. The results describe performance in the benchmark’s tested setup. The supplied source does not establish that the scores predict performance across different live systems or long-term use.

Where can developers run ThinkingBox?

The release says it is available through Hugging Face and can be run through OpenEnv using isolated MCP tool sessions.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Gewerkton Used AI Agents To Rethink Its Look On A Two-Day Timeline

Gewerkton says AI coding agents implemented its HORIZON redesign across apps and a 40-language website on Oct. 2-3, with human review and fixes.

Gaming Revolution: Atari’s Unexpected Surge Across The Globe

Atari experiences an unexpected surge in worldwide popularity, marking a major shift in the gaming industry. Details are still emerging.

#ASUS Trending In The Fediverse

ASUS has recently become a trending topic on Mastodon, with multiple accounts discussing the brand. The development highlights growing social media interest in the company.

10 Best Network Attached Storage Devices For Private Cloud Storage In 2026

Discover the 10 best network attached storage devices for private cloud in 2026, highlighting features, capabilities, and suitability for different users.