Why Agents Last Exam Scores Jumped: Computer Use, Bigger Models, Diverse RL Envs

dejavucoder · x · 2026-09-05

The author breaks down the Agents Last Exam benchmark: it is computer-use heavy and includes agentic knowledge-work tasks across 55+ domains such as CAD modeling. Better computer use, bigger models, and more diverse RL environments together explain the drastic increase in ALE scores.

Original post →

More from Models

Models channel →