Skip to content
Apixo
Blog
news· 2 min read· via Hugging Face blog

Evaluating AI Agents: Microsoft and Hugging Face Release ThinkingBox

Microsoft and Hugging Face released ThinkingBox, a benchmark evaluating AI agents on actual database states rather than generated text across 507 stateful workflows.

Evaluating AI Agents: Microsoft and Hugging Face Release ThinkingBox

Evaluating artificial intelligence agents has traditionally relied on surface-level outputs like generated text or successful tool calls. However, a joint initiative between Microsoft and Hugging Face aims to change this standard. The newly released benchmark, known as ThinkingBox, evaluates AI agents based on the actual database records and side effects they produce, testing whether models can perform reliably across repeated trials.

Created in partnership with Toloka and academic collaborators, ThinkingBox is now available on Hugging Face. The benchmark runs agents against isolated Model Context Protocol (MCP) tool sessions, grading terminal backend states rather than trusting final dialogue responses. According to the research, single-attempt successes often mask underlying reliability issues.

Understanding the Benchmark and Repeatability

ThinkingBox measures performance across 507 stateful business workflows in domains such as retail, auto insurance, travel, neobanks, and consulting. Each workflow runs 20 independent times from a clean backend state to determine consistency.

The research outlines three core metrics: pass@1 (single-attempt success), pass@20 (solving a task at least once in 20 tries), and observed 20/20 (tasks passing every single attempt). While models like Kimi-K3 show high breadth by solving many tasks at least once, they frequently lack consistency. Conversely, models like Claude Opus 5 and Claude Opus 5.5 demonstrate stronger consistency across all 20 attempts, though higher headline accuracy does not always guarantee dependability.

Cost efficiency also plays a vital role when deploying these systems. The analysis prices model usage based on undiscounted list rates to evaluate the cost per successful task and the cost per dependable task. The Pareto cost frontier highlights models like GPT-5.6 Sol, GPT-5.4, and Claude Opus 5.5 as offering strong trade-offs between cost and accuracy. Meanwhile, diagnostic signatures reveal that roughly 80% of failures stem from tool handling and error recovery rather than high-level reasoning deficits.

What it means for developers

For developers building production-grade automation, ThinkingBox highlights the danger of relying on single-run evaluations. Because an agent can execute valid tool calls while leaving incorrect database fields or unaddressed exceptions, engineering teams must test terminal backend states instead of trusting model summaries. Developers can try top AI models cheaply through one API at https://apixoai.online. Implementing rigorous error classification, minimizing unnecessary tool surfaces, and requiring human approvals for irreversible actions remain crucial steps for reliable agent deployment.

Running ThinkingBox

ThinkingBox and ThinkingBox-Bench are now accessible via the OpenEnv interface on Hugging Face. The setup requires Python 3.11+, uv, Docker, a pinned release of the thinkingbox-data repository, and Typesense for search indexing. Developers can execute individual evaluation tasks through the packaged CLI and inspect operational errors using dedicated sidecar outputs.


Source: The Agent Said It Was Done. The Database Disagreed. — Hugging Face blog. Written by the Apixo team from that report.

#ai-news#ai-agents#benchmarks#microsoft#hugging-face#developers
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading