Microsoft and Hugging Face released ThinkingBox, a benchmark that evaluates AI agents by grading their impact on isolated database states rather than just their tool calls or final responses. Across 507 stateful business workflows run 20 times each against various LLM models, the study found that while many agents produce valid tool calls, they frequently leave incorrect values, unintended extra effects, or missing required effects in the backend.
- Claude Opus 5.5 leads overall with a 67.16% pass@1 score, but only three models retain most of their single-attempt scores across 20 repeats: GPT-6 Astra (78%), Claude Opus 5.5 (71%), and Claude Opus 5 (71%).
- Kimi-K3 solves the broadest coverage of tasks (93.89% at least once) but is among the least consistent, succeeding in all 20 attempts for only 13.41% of tasks.
- GPT-5.6 Sol offers the lowest cost per successful task attempt at $0.127, while GPT-5.4 is the cheapest dependable option at $6.80 per task passed on all 20 attempts.
- Approximately four in five failures are attributed to tool handling issues, such as failing to recover from errors or empty lookups, rather than reasoning deficits.
The authors argue that pass@1 scores alone are insufficient for deployment reliability and recommend treating the 20/20 consistency rate as a design input, checking terminal state before committing changes in production.