ApprenticeBench: best human testers pass only 51% first-try — agents don't need perfection

hhsun1 · x · 2026-09-16

While building ApprenticeBench (a benchmark for computer use + continual learning on real jobs), moyicat raises a key question: if our best human testers score only 51%, why do we treat their real-world accounting work as near-flawless?

The answer is checks and balances: in real jobs multiple people review each bill, with dashboards and reconciliation catching anomalies. Human performance comes from teams and well-designed mechanisms, not one person scoring 100%.

The implication: AI agents don't need first-try perfection to be useful — they need to learn on the job, respond to corrections, and fit into existing systems. The original announcement claims Fable 5.1 and GPT-6 Astra can continually learn and surpass human professionals on this benchmark, deploying themselves without FDEs.

Original post →

More from coding & agent

coding & agent channel →