AutoResearchExam: a 24-hour-per-model benchmark tests AI research agents' long-horizon skills
sudoraohacker · x · 2026-09-10
Alex Dimakis's team released AutoResearchExam, a benchmark of open-ended ML and engineering tasks across seven research areas including model training, data curation, AI safety and interpretability.
- Each task gives an agent a CPU/GPU machine and 24 hours to improve solutions via experiments; scoring combines speed and quality.
- Unique feature: tests whether agents' improvements hold up on unseen data.
- Key finding: AI research agents often overfit while trying to improve.
- Running one model takes 24 hours; frontier head-to-head comparisons (e.g., Astra) are underway.
Related event: AutoResearchExam: Benchmarking AI Agents on 24-Hour Autonomous Research(6 posts)→
More from coding & agent
- GPT-6 Astra drives Houdini for hands-off procedural modeling with one prompt — ssh4net · 2026-09-10
- SAEScientist-Bench: can AI agents autonomously run SAE interpretability research? — CASIA · 2026-09-10
- Dev uses GPT Astra to port a 33-year-old game to HTML5 with online mode — Vintendopower · 2026-09-10
- mcp-oracle-h adds a mandatory human approval gate for irreversible agent actions via Telegram — modelcontextprotocol · 2026-09-10
- Codex already lets you group pinned sessions by dragging them into projects — vwxyzjn · 2026-09-10
- Use /goal to steer the astra coding agent back on track — TheMoonMidas · 2026-09-10