A joint blog post from Microsoft and Hugging Face describes a new agent evaluation framework, ThinkingBox, built around a simple idea: grade an AI agent by what it actually wrote to a database, not by what it said in its final reply.
The motivating example is a support agent handling a late delivery. It makes nine tool calls, reads the refund policy correctly, and closes the ticket as resolved. Two things are still wrong: the courier exception is still open, so the ticket should have been left on hold, and the customer never got an answer to her actual question. A grader checking only tool calls or the final sentence would mark this a success. The database disagrees.
ThinkingBox-Bench runs 507 stateful business workflows across retail, auto insurance, travel, neobank, and consulting domains, each repeated 20 times per model against 18 proprietary and open-weight LLMs, with every attempt starting from an identical clean backend. In a common-set ablation of 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks even though 67.24% of those failures terminated cleanly with no reported tool error. Of the failures, 77.61% had wrong field values, 43.30% produced unintended extra effects, and 25.36% were missing a required effect.
Single-attempt accuracy (pass@1) and consistency across all 20 tries (observed 20/20) tell very different stories. Claude Opus 5.5 leads overall pass@1 at 67.16%, with Kimi-K3 the strongest open-weight model at 57.37%. But Kimi-K3 solves 93.89% of tasks at least once while passing only 13.41% (68 of 507) on every single attempt; Claude Opus 5 solves fewer tasks at least once (79.09%) but passes 47.53% (241 of 507) every time. A newer model, Claude Opus 5.5, scored higher on average than Claude Opus 5 but passed the exact same number of tasks (241) on all 20 attempts, meaning the accuracy gain bought no added dependability.
The diagnostic breakdown of failures is the most actionable part: 79.9% are tool-usage errors, 10.3% are wrong state updates, 7.0% are incomplete user resolutions, and only 2.9% involve no state-changing action at all. That means most failures are retry-and-recovery problems, not reasoning problems.
ThinkingBox is now available through Hugging Face and OpenEnv, with the framework under MIT license and the benchmark data under CDLA-Permissive-2.0.
Top comments (0)