Microsoft's Thinkingbox: best model hits only 65% pass@1 on real business workflows

dair_ai · x · 2026-08-23

Microsoft released Thinkingbox, a sandbox plus benchmark for evaluating agent reliability in real business workflows:

The strongest model reaches just 65.36% pass@1 and 25.25% pass^20. Notably, many failed trials terminate cleanly with valid state-changing tool calls — watching the response or tool call tells you almost nothing about whether the task actually completed.

Related event: Microsoft Paper: Agents Succeed Often but Rarely Reliably(2 posts)→

Original post →

More from coding & agent

coding & agent channel →