First end-to-end benchmark tests AI continual learning in a real job

ysu_nlp · x · 2026-09-11

A new paper introduces the first end-to-end benchmark spanning computer use, continual learning, and long-horizon agentic capabilities, set inside a real job. The authors argue that the two main roadblocks to AGI—recursive self-improvement and continual learning—share the same spirit: AI must not only act like humans but teach itself like humans in a fast-moving, time-sensitive world. The benchmark observes how frontier models act and learn like an apprentice in a real workplace.

Related event: NeoCognition Releases ApprenticeBench, First Job-Level Continual Learning Benchmark(8 posts)→

Original post →

More from Research

Research channel →