According to a joint Microsoft and Hugging Face post, ThinkingBox—now available through Hugging Face—scores agents on the terminal backend state and side effects they leave behind. In one benchmark task, an agent handling a delayed appliance order made nine well-formed tool calls, marked the ticket resolved, and told the customer its query was resolved. The required end state was "hold" because the carrier exception was still open; the database showed the agent had not actually finished.
ThinkingBox runs 507 stateful business workflows 20 times each from identical clean backends, reporting pass@1, pass@20, and observed 20/20. In a common-set ablation covering 121,680 valid trials across 12 LLMs, 79,853 attempts failed executable checks. Among those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error; state checks nevertheless found wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36%.
On pass@1, Claude Opus 5.5 led overall at 67.16%, just ahead of Claude Opus 5 at 66.50%. Kimi-K3 was the strongest open-weights model at 57.37%, within a point of GPT-6 Astra. The post's central point is that a tool call is not an outcome and one success is not reliability: repetition and the records left behind are what separate a working agent from a plausible one.