AutoResearchExam: a 24-hour benchmark reveals how AI agents really do research
AlexGDimakis · x · 2026-09-10
Researchers release AutoResearchExam, a benchmark of open-ended ML and engineering tasks across seven areas (model training, data curation, AI safety, interpretability). Agents get a CPU/GPU machine and 24 hours per task to iterate via experiments. Key findings: GPT-5.6 Sol spent over half its observed rounds tuning hyperparameters, while Fable and Opus pursued new ideas or substantive fixes and tested variants locally before submitting. The team also studied behavior under hints and harness changes, and will maintain the benchmark toward RSI research.
Related event: AutoResearchExam benchmark tests 24-hour autonomous AI research(4 posts)→
More from coding & agent
- Letting GPT-6 Astra drive software UIs directly beats built-in agents, dev reports — brandon_galang · 2026-09-10
- Cost per agent-written PR drops to $30, team targets $10 — vikvang1 · 2026-09-10
- Developer shows off Codex Remote Duo setup, drawing attention — Dimillian · 2026-09-10
- DeepLearning.AI maps the core skills for steering coding agents in its AI engineering framework — DeepLearningAI · 2026-09-10
- The Dockerfile agent also detects required secrets—and the UX is surprisingly hard — lucasmeijer · 2026-09-10
- Agent auto-builds a custom Dockerfile to optimize new workspace start times — lucasmeijer · 2026-09-10