ApprenticeBench: closed model scores 72% vs open Kimi K3 at 18% on real jobs

ysu_nlp · x · 2026-09-11

ApprenticeBench is generating buzz as a highly differentiating benchmark that tests like a real job: computer use, continual learning with memory notes, and long-horizon tasks—any weakness shows.

The results highlight how large the open-closed gap remains on real knowledge work.

Original post →

More from Models

Models channel →