ApprenticeBench claims strongest model separation: 72% vs 18% job completion for near-tied models

ysu_nlp · x · 2026-09-11

Quoting the ApprenticeBench launch thread, the author argues it may be the most discriminating agent benchmark yet, showing a step change for Fable 5.1 and GPT-6 Astra.

Related event: ApprenticeBench Tests Agents Learning Real Jobs, Closed Models Lead by Wide Margin(10 posts)→

Original post →

More from AGI Musings

AGI Musings channel →