HANDBOOK.md benchmarks whether agents obey long SOPs over extended tool use
dair_ai · x · 2026-07-29
Surge AI’s new benchmark, HANDBOOK.md, tests whether agents actually follow long-standing instructions over extended tool use, not just whether they reach the right answer.
- The benchmark is designed for enterprise-style settings where system prompts, policy files, or skills documents are supposed to bind behavior over time.
- It includes 65 agentic tasks across five domains and ten fictional companies, each in a self-contained environment with file workspace plus mock email, chat, calendar, issue tracking, and commerce tools exposed via MCP.
- Tasks use 20- to 124-page handbooks, mutate one of ten base SOPs, and are graded with deterministic programmatic rubrics checking both required and forbidden actions.
- The paper reports that under strict grading, the best of 31 evaluated model configurations passes only 36.2% of tasks, with most frontier configs below 25%.
More from coding & agent
- ChatGPT vs Claude Browser Agents: Login Persistence Makes the Difference — dkundel · 2026-07-30
- The Worst Enterprise AI Retrieval Bug: Confident but Incomplete Answers — rohanpaul_ai · 2026-07-30
- Pulse brings local network logging and request inspection to Apple apps — tom_doerr · 2026-07-30
- Rust terminal diff viewer adds AI commit messages and PR review — iamsahaj_xyz · 2026-07-30
- OpenSquilla v0.5.1 splits model access from routing and claims big cost savings — hey_abusiddik · 2026-07-30
- Google ADK article argues tool-call interceptors are the missing safety layer for agents — fhinkel · 2026-07-30