🔍 Read the full analysis: What Happens When An AI Agent’s Work Doesn’t Match The Database? on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Microsoft and Hugging Face have made ThinkingBox available through Hugging Face. The benchmark tests 507 business workflows, checking backend records and side effects across 20 runs per task; its authors report frequent failures in their tested setup, even when agents do not report a final tool error.
Microsoft and Hugging Face have made ThinkingBox, a benchmark for AI agents, available through Hugging Face. It tests whether agents leave business systems in the required state—not just whether they issue valid tool calls or produce a plausible reply—across 507 workflows repeated 20 times per task, addressing the kind of reliability gap highlighted in the original analysis.
ThinkingBox runs agents in isolated sessions using MCP tools, then checks the backend’s final state and any side effects against executable requirements. The tasks cover retail, auto insurance, travel, neobanking and consulting. Repeating each workflow offers a measure of whether an agent can complete it consistently in the benchmark setup, rather than only succeeding in one attempt. This focus on agents acting through tools also connects to platforms for agents and apps.
The release describes a retail support task involving a delayed $745 appliance order. The agent investigates the order, opens a ticket and records a timeline. The customer is not eligible for late-delivery compensation under the policy the agent checked, but the carrier exception remains open and the ticket is supposed to stay on hold. The agent instead marks it solved and replies without addressing the customer’s underlying question. The executable check fails because the ticket status is solved rather than hold.
In a common-set analysis of 121,680 valid trials across 12 models, the authors report 79,853 attempts failed executable checks. Of those failures, 67.24% ended without a final tool error despite an agent having invoked a state-changing tool. The authors also report wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%. These categories overlap, so a failed attempt could appear in more than one.
The release reports an overall pass@1 of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, identified as the strongest open-weight model in the table. Pass@1 is the share of attempts that succeed. The authors also define pass@20 as whether a task succeeds at least once in 20 runs, and observed 20/20 as passing every recorded run. They say Kimi-K3 scored within one point of GPT-6 Astra, but the supplied material does not include uncertainty estimates for that comparison.
Why Database Outcomes Change Agent Evaluation
For companies using agents on support tickets, refunds, claims or bookings, a polished response is not proof that the requested work was completed. An agent can make tool calls without a final error and still leave the wrong value in a record, close a case too soon, or create an unintended side effect. Checking backend state evaluates the result that matters to the workflow, not only the conversation that precedes it.
The repeated runs address a separate question: how consistently an agent performs. A single success does not establish that a task will work reliably on subsequent attempts. Passing all 20 observed runs is stronger evidence within this benchmark than passing once, but it remains a limited set of trials—not a guarantee of dependable performance in a live organization.
ThinkingBox can give developers and evaluators a way to compare models and locate workflow failures under a shared test setup. Its scores do not, by themselves, establish how a model will perform with a company’s records, policies, integrations or unusual customer requests. Organizations would still need to test their own workflows and inspect failures as well as successful runs.
As an affiliate, we earn on qualifying purchases.
From Tool Calls to Final System State
The benchmark authors frame ThinkingBox around a distinction between actions and outcomes. A tool call may be valid and a response may sound appropriate, but neither necessarily shows that the system record now meets the task’s requirements. The benchmark’s executable checks inspect that final record and related side effects.
According to the release, each workflow is run from a clean backend, and the benchmark is based on the authors’ paper. It can be run through OpenEnv. The supplied material does not give a publication date or the full paper’s evaluation details, so readers cannot assess every task specification, model configuration or comparison from the excerpt alone.
“A tool call is not an outcome.”
— Microsoft and Hugging Face, in the ThinkingBox release
database integrity monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Benchmark Results
The reported failure rates and model scores are the authors’ results in the tested setup. The supplied release material does not include the full evaluation table, uncertainty estimates, complete model configurations or all task specifications. It therefore does not establish how meaningful small score differences are or whether the rankings would hold under other conditions.
It is also unclear how well the 507 workflows represent the range of tasks, changing records and edge cases in live business systems. Twenty runs provide a bounded observation; the material does not show that passing all 20 predicts long-term reliability. The supplied source gives no publication date, independent replication results or evidence that benchmark scores directly predict performance at a particular organization.
As an affiliate, we earn on qualifying purchases.
Testing Agent Workflows Beyond the Benchmark
The release says developers can run ThinkingBox through OpenEnv with isolated MCP tool sessions, inspect the tasks and compare agent outcomes against executable checks. The supplied material does not name a future release date or another planned milestone.
For organizations considering agent deployment, the immediate practical step is to test workflows against their own business rules and backend requirements. That means checking not only whether the agent responds or completes a tool call, but whether it leaves the correct records and avoids unintended changes. Whether ThinkingBox’s results predict performance in any particular company remains to be tested.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does ThinkingBox test?
It checks whether an AI agent leaves a business system in the required final state, including the relevant records and side effects, rather than relying only on tool-call validity or response quality.
How large is the benchmark?
The release describes 507 workflows, each repeated 20 times per task. The reported common-set analysis covers 121,680 valid trials across 12 models.
What did the authors report about failures?
They report that 79,853 attempts in the common-set analysis failed executable checks. Among those failures, 67.24% ended without a final tool error despite a state-changing tool call. The failure categories for wrong values, extra effects and missing effects overlap.
Do the reported scores show how agents will perform at a company?
No. The scores describe performance in the benchmark’s tested setup. They do not establish performance across every live business system, company policy or unusual workflow.
Where can developers run ThinkingBox?
The release says it is available through Hugging Face and can be run through OpenEnv using isolated MCP tool sessions.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
