Microsoft's Thinkingbox: best model hits only 65% pass@1 on real business workflows
dair_ai · x · 2026-08-23
Microsoft released Thinkingbox, a sandbox plus benchmark for evaluating agent reliability in real business workflows:
- The sandbox provides isolated MCP-compatible tool sessions.
- The benchmark spans 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank IT, and consulting support.
- Grading ignores what the agent says and inspects the backend state it leaves behind: executable checks accept valid trajectories and reject wrong, missing, or extra effects — collateral damage counts against you.
The strongest model reaches just 65.36% pass@1 and 25.25% pass^20. Notably, many failed trials terminate cleanly with valid state-changing tool calls — watching the response or tool call tells you almost nothing about whether the task actually completed.
Related event: Microsoft Paper: Agents Succeed Often but Rarely Reliably(2 posts)→
More from coding & agent
- Agentic coding accessibility will reshape understanding of software complexity — pixlpa · 2026-08-24
- Devin Agent bypasses Slack block by finding emails in git logs — sandylikesfrogs · 2026-08-24
- Developer habits shift: Agents become collaborators from simple tools — latticecut · 2026-08-24
- Dev bottleneck shifts from writing to reading code: exe.dev co-founder — thursdai_pod · 2026-08-24
- The biggest AI mistake: trying to reinvent the wheel instead of using tools — Tired40s · 2026-08-24
- DeepPaperNote turns research papers into Obsidian notes — tom_doerr · 2026-08-24