Why Agents Last Exam Scores Jumped: Computer Use, Bigger Models, Diverse RL Envs
dejavucoder · x · 2026-09-05
The author breaks down the Agents Last Exam benchmark: it is computer-use heavy and includes agentic knowledge-work tasks across 55+ domains such as CAD modeling. Better computer use, bigger models, and more diverse RL environments together explain the drastic increase in ALE scores.
More from Models
- Report: Claude may have cracked the Navier–Stokes Millennium Problem, pending expert scrutiny — MoonL88537 · 2026-09-06
- OpenCUA author: the 450-step Paint demo we rejected 2 years ago is now doable by agents — andersonbcdefg · 2026-09-06
- Yacine: Astra Is Now Fast Enough for Real-Time Iterative Work — yacineMTB · 2026-09-06
- Chollet: OpenAI's New Model Scores Near-Perfect on ARC-AGI-3 for the First Time — fchollet · 2026-09-06
- Xbench turns Twitter into an AI eval, tracking real sentiment and model switches — thedealdirector · 2026-09-06
- Astra refuses to score AI governance stories, researcher flags model behavior — sethlazar · 2026-09-06