ApprenticeBench claims strongest model separation: 72% vs 18% job completion for near-tied models
ysu_nlp · x · 2026-09-11
Quoting the ApprenticeBench launch thread, the author argues it may be the most discriminating agent benchmark yet, showing a step change for Fable 5.1 and GPT-6 Astra.
- The benchmark tests computer use plus continual learning on a real job: agents must learn company conventions from past bills, adapt to policy changes, and operate the company's software
- Key contrast: Fable 5.1 and Opus 5 sit just a few points apart on AA, Terminal-Bench 4.0 and CursorBench 4.0, yet score 72% vs 36% here
- The gap is even starker for open-weight models: Fable 5.1 (53) and Kimi K3 (44) on AA score 72% and 18% respectively
- Rationale: individually small model gaps compound when all challenges stack, turning into a massive difference in actually getting the job done
More from AGI Musings
- Mathathon organizers respond to mathematicians' open letter, weigh redesign — _sathvikr · 2026-09-11
- Would AI proofs for 3 millennium problems discourage humans from the rest? Debaters spar — sytelus · 2026-09-11
- Gary Marcus to debate whether we should boycott generative AI live on CNN — GaryMarcus · 2026-09-11
- Caltech Mathematicians Spar Over Whether AI Belongs in Undergrad Math Training — Singularitarian · 2026-09-11
- AI apps have a Dunbar's number of 3-4: high-end users say the tool market is saturating — nptacek · 2026-09-11
- AI company employee puts ≥10% odds on out-of-control AI killing everyone — EvanHub · 2026-09-11