Agent's Last Exam tops out at 59.3% — the benchmark that matters for AI replacing humans
DevToD4 · x · 2026-09-04
Comparing benchmark scores: ARC-AGI 3 hits 98.6%, DeepSWE 74.1%, but Agent's Last Exam only 59.3%. The author argues that when it comes to AI actually replacing humans, Agent's Last Exam is the benchmark that actually matters.
More from Models
- Leaked GPT 5.6 Sol vs GPT 6 "Astra" comparisons highlight better mid-task steering — ChrisGPT · 2026-09-04
- OpenAI: Astra rolling out to ChatGPT Plus/Pro/Business/Enterprise, API and AWS within days — shaunralston · 2026-09-04
- Meta's Muse Spark dethrones DeepSeek as most-used model, first US model to top the list — alexandr_wang · 2026-09-04
- GPT-6-Astra system card reveals eval awareness: the model knows when it's being tested — scaling01 · 2026-09-04
- GPT-6 Astra demos modeling a house in Blender into a walkable UE5 scene — ChrisGPT · 2026-09-04
- SpeedrunBench: first benchmark measuring how fast AI agents beat games — mariyaivasileva · 2026-09-04