Microsoft and Hugging Face detail ThinkingBox agent reliability results
A new joint analysis says many agents complete tool calls yet leave incorrect business records, and repeated success varies sharply by model.
Checking the result an agent leaves behind
Microsoft and Hugging Face published a joint ThinkingBox analysis on October 3 focused on whether AI agents leave business systems in the requested state. The authors describe a customer-support example in which an agent makes valid tool calls but closes a ticket that should remain on hold. Their benchmark checks final database values and side effects, then repeats each of 507 stateful workflow tasks 20 times from a clean starting state. The paper and public dataset predate this blog post; the October publication is a new explanation of the results, not the first release of the benchmark.
The authors report that 79,853 of 121,680 valid trials failed executable checks in one 12-model analysis. Of those failures, 67.24% still finished without a final tool error after invoking a state-changing tool. The post also contrasts single-attempt scores with tasks completed correctly on all 20 runs. Those figures come from the researchers’ controlled workflows and should not be read as measured failure rates for deployed customer agents. The practical point is narrower: a fluent final answer or a successful tool call does not establish that the intended record was changed correctly.