ApprenticeBench: closed model scores 72% vs open Kimi K3 at 18% on real jobs
ysu_nlp · x · 2026-09-11
ApprenticeBench is generating buzz as a highly differentiating benchmark that tests like a real job: computer use, continual learning with memory notes, and long-horizon tasks—any weakness shows.
- Closed Fable 5.1 scores 72% at $18.23 per task; open Kimi K3 manages only 18% while costing more ($25.83).
- Frontier models are already more accurate than human professionals, but at roughly 250% of the cost.
- Open models aren't necessarily cheaper: struggling models burn tokens, pushing total cost above closed rivals.
The results highlight how large the open-closed gap remains on real knowledge work.
More from Models
- DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture — ccerrato147 · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions — Scobleizer · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- Claims resurface that Moonshot's Kimi distilled from Claude raw CoTs — xuanalogue · 2026-09-11