Auditing Agents' Last Exam: Porting the ALE Linux CLI subset to Verifiers v1 reveals widespread data flaws

dejavucoder · x · 2026-10-11

The author ported the Linux CLI subset of Agents' Last Exam (ALE), a frontier agent benchmark by Dawn Song's group at UC Berkeley, to Verifiers v1 — and found serious data issues along the way:

This led to building ALE-Gold, a cleaned-up subset, plus full eval runs comparing Sol (xhigh) vs Luna (xhigh), failure-mode analysis, and per-domain breakdowns. Bottom line: a benchmark built with 300+ experts across 100+ institutions still has substantial engineering room for improvement. A detailed, hands-on work log worth reading for anyone doing evals.

Original post →

More from Research

Research channel →