AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Happens When An AI Agent’s Work Doesn’t Match The Database? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Microsoft and Hugging Face have made ThinkingBox available through Hugging Face. The benchmark tests 507 business workflows, checking backend records and side effects across 20 runs per task; its authors report frequent failures in their tested setup, even when agents do not report a final tool error.

Microsoft and Hugging Face have made ThinkingBox, a benchmark for AI agents, available through Hugging Face. It tests whether agents leave business systems in the required state—not just whether they issue valid tool calls or produce a plausible reply—across 507 workflows repeated 20 times per task, addressing the kind of reliability gap highlighted in the original analysis.

ThinkingBox runs agents in isolated sessions using MCP tools, then checks the backend’s final state and any side effects against executable requirements. The tasks cover retail, auto insurance, travel, neobanking and consulting. Repeating each workflow offers a measure of whether an agent can complete it consistently in the benchmark setup, rather than only succeeding in one attempt. This focus on agents acting through tools also connects to platforms for agents and apps.

The release describes a retail support task involving a delayed $745 appliance order. The agent investigates the order, opens a ticket and records a timeline. The customer is not eligible for late-delivery compensation under the policy the agent checked, but the carrier exception remains open and the ticket is supposed to stay on hold. The agent instead marks it solved and replies without addressing the customer’s underlying question. The executable check fails because the ticket status is solved rather than hold.

In a common-set analysis of 121,680 valid trials across 12 models, the authors report 79,853 attempts failed executable checks. Of those failures, 67.24% ended without a final tool error despite an agent having invoked a state-changing tool. The authors also report wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%. These categories overlap, so a failed attempt could appear in more than one.

The release reports an overall pass@1 of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, identified as the strongest open-weight model in the table. Pass@1 is the share of attempts that succeed. The authors also define pass@20 as whether a task succeeds at least once in 20 runs, and observed 20/20 as passing every recorded run. They say Kimi-K3 scored within one point of GPT-6 Astra, but the supplied material does not include uncertainty estimates for that comparison.

At a glance
reportWhen: Availability announced in the supplied…
The developmentMicrosoft and Hugging Face have made ThinkingBox, a benchmark that checks AI agents’ resulting database states and workflow side effects, available through Hugging Face.
At a glance
announcementWhen: Now available through Hugging Face; the…
The developmentMicrosoft and Hugging Face released ThinkingBox through Hugging Face, a benchmark for evaluating AI agents by their backend changes across repeated workflow trials.

Why Database Outcomes Change Agent Evaluation

For companies using agents on support tickets, refunds, claims or bookings, a polished response is not proof that the requested work was completed. An agent can make tool calls without a final error and still leave the wrong value in a record, close a case too soon, or create an unintended side effect. Checking backend state evaluates the result that matters to the workflow, not only the conversation that precedes it.

The repeated runs address a separate question: how consistently an agent performs. A single success does not establish that a task will work reliably on subsequent attempts. Passing all 20 observed runs is stronger evidence within this benchmark than passing once, but it remains a limited set of trials—not a guarantee of dependable performance in a live organization.

ThinkingBox can give developers and evaluators a way to compare models and locate workflow failures under a shared test setup. Its scores do not, by themselves, establish how a model will perform with a company’s records, policies, integrations or unusual customer requests. Organizations would still need to test their own workflows and inspect failures as well as successful runs.

Amazon

AI workflow testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Tool Calls to Final System State

The benchmark authors frame ThinkingBox around a distinction between actions and outcomes. A tool call may be valid and a response may sound appropriate, but neither necessarily shows that the system record now meets the task’s requirements. The benchmark’s executable checks inspect that final record and related side effects.

According to the release, each workflow is run from a clean backend, and the benchmark is based on the authors’ paper. It can be run through OpenEnv. The supplied material does not give a publication date or the full paper’s evaluation details, so readers cannot assess every task specification, model configuration or comparison from the excerpt alone.

“A tool call is not an outcome.”

— Microsoft and Hugging Face, in the ThinkingBox release

Amazon

database integrity monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Benchmark Results

The reported failure rates and model scores are the authors’ results in the tested setup. The supplied release material does not include the full evaluation table, uncertainty estimates, complete model configurations or all task specifications. It therefore does not establish how meaningful small score differences are or whether the rankings would hold under other conditions.

It is also unclear how well the 507 workflows represent the range of tasks, changing records and edge cases in live business systems. Twenty runs provide a bounded observation; the material does not show that passing all 20 predicts long-term reliability. The supplied source gives no publication date, independent replication results or evidence that benchmark scores directly predict performance at a particular organization.

Amazon

business process automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing Agent Workflows Beyond the Benchmark

The release says developers can run ThinkingBox through OpenEnv with isolated MCP tool sessions, inspect the tasks and compare agent outcomes against executable checks. The supplied material does not name a future release date or another planned milestone.

For organizations considering agent deployment, the immediate practical step is to test workflows against their own business rules and backend requirements. That means checking not only whether the agent responds or completes a tool call, but whether it leaves the correct records and avoids unintended changes. Whether ThinkingBox’s results predict performance in any particular company remains to be tested.

Amazon

AI agent reliability testing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does ThinkingBox test?

It checks whether an AI agent leaves a business system in the required final state, including the relevant records and side effects, rather than relying only on tool-call validity or response quality.

How large is the benchmark?

The release describes 507 workflows, each repeated 20 times per task. The reported common-set analysis covers 121,680 valid trials across 12 models.

What did the authors report about failures?

They report that 79,853 attempts in the common-set analysis failed executable checks. Among those failures, 67.24% ended without a final tool error despite a state-changing tool call. The failure categories for wrong values, extra effects and missing effects overlap.

Do the reported scores show how agents will perform at a company?

No. The scores describe performance in the benchmark’s tested setup. They do not establish performance across every live business system, company policy or unusual workflow.

Where can developers run ThinkingBox?

The release says it is available through Hugging Face and can be run through OpenEnv using isolated MCP tool sessions.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

NYT Connections Answers for July 3, 2026

The New York Times has published the answers for the July 3, 2026, NYT Connections puzzle, with official solutions now available online.

Interview with Mitchell Hashimoto about Ghostty and Zig

Mitchell Hashimoto shares insights on Ghostty and Zig, highlighting their roles in modern infrastructure and development. Key details from the recent interview.

Augmented Reality (AR) on Phones: What It Is and How It’s Used

AIThis post was created with the assistance of artificial intelligence (AI).Augmented reality…

Optimizing 350M AI Models For Superior Structured Results With Limited Training Steps

Liquid AI publicly releases a low-cost fine-tuning method for 350M models, boosting schema compliance on the IFStruct benchmark from 22.6% to 29.7% with minimal resources.