Agent's Last Exam tops out at 59.3% — the benchmark that matters for AI replacing humans

DevToD4 · x · 2026-09-04

Comparing benchmark scores: ARC-AGI 3 hits 98.6%, DeepSWE 74.1%, but Agent's Last Exam only 59.3%. The author argues that when it comes to AI actually replacing humans, Agent's Last Exam is the benchmark that actually matters.

Original post →

More from Models

Models channel →