Porting Agents' Last Exam's Linux CLI subset surfaced benchmark defects, yielding ALE-Gold
dejavucoder · x · 2026-09-18
During Prime Residency, the author ported the Linux CLI subset of Agents' Last Exam (a frontier benchmark from UC Berkeley's Dawn Song group) to Verifiers v1—and digging into the data revealed numerous issues, leading to a corrected ALE-Gold.
- Porting challenges: minimizing rollout time, missing dependency specs, and defects in freshly released Verifiers v1 (since fixed)
- Data problems found: underspecified task contracts, evaluator defects, answer leakage and network issues, broken inputs, missing/corrupt reference data
- Released full eval runs comparing Sol (xhigh) vs Luna (xhigh), selected task comparisons, failure modes, and per-domain performance
- Conclusion: engineering still has significant room for improvement
Related event: Researcher Reproduces Frontier Benchmark ALE, Finds Data Issues(3 posts)→
More from Research
- NVIDIA tutorial: memory-driven self-model agent hits 90.9% vs 82.8% RAG baseline — dl_weekly · 2026-09-18
- New paper: Tri-Metric Router cuts long-context RAG OOM failures to 0% on a T4 — chaumian · 2026-09-18
- Data Lab Argues Truly General Synthetic Data Matters More Than Architecture — Paimaamu · 2026-09-18
- 3DV 2027 opens Nectar Track call for papers, deadline Feb 15, 2027 — ftm_guney · 2026-09-18
- New Tool Lets You Search and Skim Anthropic's Released Mythos Transcript by Flagged Behaviors — round · 2026-09-18
- Experiment calibrates 48 attention heads down to 12, swapping the rest for band-diagonal sparse attention — ostrisai · 2026-09-18