ApprenticeBench: best human testers pass only 51% first-try — agents don't need perfection
hhsun1 · x · 2026-09-16
While building ApprenticeBench (a benchmark for computer use + continual learning on real jobs), moyicat raises a key question: if our best human testers score only 51%, why do we treat their real-world accounting work as near-flawless?
The answer is checks and balances: in real jobs multiple people review each bill, with dashboards and reconciliation catching anomalies. Human performance comes from teams and well-designed mechanisms, not one person scoring 100%.
The implication: AI agents don't need first-try perfection to be useful — they need to learn on the job, respond to corrections, and fit into existing systems. The original announcement claims Fable 5.1 and GPT-6 Astra can continually learn and surpass human professionals on this benchmark, deploying themselves without FDEs.
More from coding & agent
- Primitive rebuilds email infrastructure so AI agents get their own inbox — Madisonkanna · 2026-09-16
- When Slack API Failed His Agent, Codex Fell Back to Computer Use — and It Worked — Yuchenj_UW · 2026-09-16
- Building a shared kanban workspace for a long-running agent with WebMCP — prd_008 · 2026-09-16
- Poolday raises $11M for a video agent behind 100M+ edits, fully layer-editable output — kimmonismus · 2026-09-16
- Taste launches Brand API to keep AI agents on-brand, with free tier and MCP access — sarahcat21 · 2026-09-16
- Scale AI frames root-cause attribution of long-horizon agent failures as a continual search problem — ScaleAI · 2026-09-16