ApprenticeBench claims AI agents surpass humans on real jobs; Liang Wenfeng calls it 'prompt engineering bullshit'

ysu_nlp · x · 2026-09-16

NeoCognition introduced ApprenticeBench, testing computer use plus continual learning on real jobs, claiming Fable 5.1 and GPT-6 Astra can learn on the job and surpass human professionals, deploying themselves with no forward-deployed engineers.

DeepSeek's Liang Wenfeng pushed back: "that's not continual learning, this is prompt engineering bullshit." TeortaxesTex countered that with superhuman priors, retrieval over infinite databases, and fast thinking, parametric updates may not even be needed for practical purposes.

Original post →

More from AGI Musings

AGI Musings channel →