AutoLab benchmark shows frontier models win long-horizon tasks by persisting, not guessing
rohanpaul_ai · x · 2026-07-22
AutoLab benchmark reveals persistence matters more than first guesses
The paper introduces AutoLab, a benchmark of 36 long-horizon tasks where each agent starts from a working but weak codebase and must improve it under a fixed time budget.
- The tasks span system speedups, puzzles, model development, and CUDA kernel work.
- The authors tested 17 frontier models and found that success depended less on the quality of the first idea and more on whether the agent kept testing, iterating, and using feedback well.
- Claude Opus 4.6 topped the benchmark not by always choosing the right move immediately, but by benchmarking repeatedly and folding empirical results into later attempts.
- Several other frontier models failed by either quitting early with time left or running out the clock before producing a useful submission.
The paper argues that current research agents still struggle with sustained, iterative work and that long-horizon persistence is a key capability gap.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11