AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Microsoft and Hugging Face have made ThinkingBox available through Hugging Face, a benchmark that evaluates AI agents by checking the backend records and side effects they leave behind. Its authors report results across 507 business workflows, with each task repeated 20 times, and say many failed attempts ended cleanly even though the resulting database state was wrong.

Microsoft and Hugging Face have made ThinkingBox available through Hugging Face, a benchmark that checks whether AI agents leave business databases in the required state. The release reports tests across 507 workflows, each run 20 times, targeting a practical gap: an agent can sound confident and finish its tool calls while failing to complete the requested work.

ThinkingBox runs agents through isolated Model Context Protocol (MCP) tool sessions and evaluates the resulting backend state and side effects. The benchmark’s authors argue that valid tool calls and plausible final messages are only indirect signs of success; the saved records show whether the task was actually completed as specified.

In a common-set comparison across 121,680 valid trials involving 12 language models, the authors say 79,853 attempts failed executable checks. Among those failed attempts, 67.24% still ended without a tool error after making a state-changing call. The checks found incorrect field values in 77.61% of failed attempts, unintended extra effects in 43.30%, and missing required effects in 25.36%. These categories overlap, according to the report.

The benchmark also measures repeatability. Each workflow is run 20 independent times from an identical clean backend. Its authors report pass@1, the share of individual attempts that succeed; pass@20, whether a task succeeds at least once in 20 attempts; and an observed 20/20 count, the number of tasks that passed every recorded run. They caution that a single successful attempt does not show that an agent will reliably repeat the result.

At a glance
announcementWhen: Announced in the Microsoft and Hugging…
The developmentMicrosoft and Hugging Face released ThinkingBox, a benchmark that tests AI agents against backend state and repeated task performance rather than relying on their tool calls or final replies.

Why Backend State Changes Matter

For businesses using agents to handle support, insurance, travel or financial workflows, a polished response is not the same as a completed transaction. A ticket may be closed prematurely, a refund may be recorded incorrectly, or an extra change may be made without a tool reporting an error. Those discrepancies can leave customers waiting and make automated work harder to audit.

ThinkingBox’s focus on persisted state gives developers and evaluators a way to test for those failures directly. Its repeated runs also distinguish a system that can succeed once from one that does so consistently under the same task conditions. The reported results suggest that clean execution traces can hide incorrect outcomes, though they do not by themselves establish how often deployed agents cause real-world harm.

The benchmark is an evaluation tool, not a guarantee that an agent will perform reliably in production. Its findings apply to the tested tasks, models and sandboxed environments described by its authors. Organizations would still need to test their own workflows, data and safeguards before relying on an agent for consequential actions.

Amazon

AI backend database testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How ThinkingBox Scores Agent Work

The report illustrates the benchmark with a customer-support scenario involving a $745 appliance delivery delayed 15 days by a courier exception. The agent checks the order, tracking, customer profile and refund policy, then opens and documents a ticket. The account does not qualify for late-delivery compensation, according to the example.

But the agent closes the ticket as solved and gives the customer no substantive answer. The required ticket status was hold, pending resolution, because the carrier exception remained open. The benchmark’s executable check flags the mismatch in the database even though the agent made nine tool calls and finished without an apparent tool failure. The authors identify the example as adapted from a specific retail task in the benchmark and point to a full trace in their paper.

The overall table in the post reports task-weighted pass@1 scores across five domains: retail, auto insurance, travel, neobank and consulting. The authors list Claude Opus 5.5 at 67.16% overall and Kimi-K3 at 57.37% among open-weight models. Those figures describe single-attempt performance on this benchmark, not a universal ranking of model quality.

“A tool call is not an outcome.”

— Microsoft and Hugging Face, in the ThinkingBox announcement

Amazon

AI workflow validation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Results

The supplied announcement does not state a publication date, and its reported scores come from the authors’ benchmark rather than an independent evaluation described here. The available material also does not establish whether the benchmark’s 507 workflows represent the full range of real business operations or how closely the isolated tool sessions match production systems.

The reported failure percentages describe failed attempts in a particular common-set comparison; they should not be read as rates of errors across all deployed AI agents. The report says the error categories overlap, and the supplied information does not provide enough detail to infer how often any one failure would affect customers or cause financial loss. It also does not specify, in the material provided, the full model-version and testing configuration for every score.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Running the Benchmark Through OpenEnv

The authors say readers can run ThinkingBox through OpenEnv and inspect the benchmark tasks and checks themselves. They present the release as a way to test agents against executable requirements, including the final state a task should produce, rather than judging responses by appearance alone.

Further assessment will depend on how developers use the benchmark, whether independent teams reproduce its findings, and how well its tasks reflect the systems they intend to automate. The announcement does not specify a next release date or a schedule for additional results. For now, its concrete next step is public access to the benchmark and the opportunity for others to run and scrutinize the tests.

Amazon

AI agent reliability testing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is ThinkingBox?

ThinkingBox is an AI-agent benchmark from Microsoft and Hugging Face. It evaluates whether an agent leaves the backend records and side effects required by a workflow.

Why check the database instead of the agent’s reply?

A reply can sound correct even when the task is unfinished or a record is wrong. ThinkingBox checks the terminal backend state to see whether the requested outcome was actually saved.

What does the 20-run test measure?

Each task starts from a clean backend and is attempted 20 times. The repeated trials show whether a task succeeds once, at least once in 20 attempts, or on all 20 recorded runs.

What did the authors report about failed attempts?

In a comparison covering 121,680 trials across 12 models, the authors report 79,853 failures on executable checks. Of those failures, 67.24% ended cleanly after a state-changing tool call, while the checks found issues including wrong values, unintended effects and missing required effects. The categories overlap.

Can the reported scores predict performance in a company’s systems?

Not on their own. The scores reflect the tested benchmark tasks and environments. A company would need to evaluate its own workflows and production conditions before relying on an agent.

Source: rss

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Ultimate Guide To Fine-tuning 350M AI Models For Improved Output Consistency

Liquid AI releases an open-source recipe to improve structured output compliance of the LFM2.5-350M model using GRPO, boosting benchmark scores on a free GPU setup.

Anthropic’s New Auto Mode For Claude: What It Means For AI Enthusiasts

Anthropic announces auto mode will become Claude’s default starting August 14, but details on operation and affected products remain unclear.

Clef: Open-source Decision Models, And New RL Fine-tuning Platform

Cloudflare has released Clef and Clef-flash on Workers AI and Hugging Face, and introduced an RL platform for customer fine-tuning.

AI-Powered Automation Software: A Labor Day sales Guide

Discover how AI-powered automation software boosts efficiency, reduces manual work, and transforms business processes—learn what you need to know now.