Auditing Agents' Last Exam: Porting the ALE Linux CLI subset to Verifiers v1 reveals widespread data flaws
dejavucoder · x · 2026-10-11
The author ported the Linux CLI subset of Agents' Last Exam (ALE), a frontier agent benchmark by Dawn Song's group at UC Berkeley, to Verifiers v1 — and found serious data issues along the way:
- Underspecified task contracts and evaluator defects
- Answer leakage and network-related failures
- Broken inputs and missing dependency specs
- Missing or corrupt reference data
This led to building ALE-Gold, a cleaned-up subset, plus full eval runs comparing Sol (xhigh) vs Luna (xhigh), failure-mode analysis, and per-domain breakdowns. Bottom line: a benchmark built with 300+ experts across 100+ institutions still has substantial engineering room for improvement. A detailed, hands-on work log worth reading for anyone doing evals.
More from Research
- Nature journal paper argues newsroom AI emotional optimization needs behavioral evaluation — lulzxdxdxd · 2026-10-11
- tangermeme encodes all 455 hg38 contigs in 1.5 seconds after auto-optimize speedups — jmschreiber91 · 2026-10-11
- tangermeme's extract_loci now 10x faster: ATAC-seq locus loading drops from 3.64s to 0.35s — jmschreiber91 · 2026-10-11
- auto-optimize speeds up genomics tool fimo ~20x with bit-identical results — jmschreiber91 · 2026-10-11
- Pointing auto-optimize at genomics code: tomtom made 7.8x faster, bit-identical output — jmschreiber91 · 2026-10-11
- 13-line lean proof of Catalan's constant irrationality; OpenAI proof controversies blamed on formalization conventions — ctjlewis · 2026-10-11