ApprenticeBench claims AI agents surpass humans on real jobs; Liang Wenfeng calls it 'prompt engineering bullshit'
ysu_nlp · x · 2026-09-16
NeoCognition introduced ApprenticeBench, testing computer use plus continual learning on real jobs, claiming Fable 5.1 and GPT-6 Astra can learn on the job and surpass human professionals, deploying themselves with no forward-deployed engineers.
DeepSeek's Liang Wenfeng pushed back: "that's not continual learning, this is prompt engineering bullshit." TeortaxesTex countered that with superhuman priors, retrieval over infinite databases, and fast thinking, parametric updates may not even be needed for practical purposes.
More from AGI Musings
- Interpretability researcher lists top open problems in decoding model activations — wesg52 · 2026-09-16
- Uncle Bob dissects the doomer debate trick of inserting nonexistent tech into doom equations — ylecun · 2026-09-16
- Catastrophe risk modeler: Tohoku tsunami shows historical data misleads on AI risk — davidmanheim · 2026-09-16
- Agent monitoring is a precondition for safe harnesses: kill CLI, perfect sandbox, or trace surveillance — scottleibrand · 2026-09-16
- Faked alignment plus recursive self-improvement: firms can't tell real alignment from theater — birchlse · 2026-09-16
- Altman warns of open-source cyberattack wave; Byrnes backs selling frontier defense — beffjezos · 2026-09-16