AutoResearchExam: a 24-hour benchmark reveals how AI agents really do research

AlexGDimakis · x · 2026-09-10

Researchers release AutoResearchExam, a benchmark of open-ended ML and engineering tasks across seven areas (model training, data curation, AI safety, interpretability). Agents get a CPU/GPU machine and 24 hours per task to iterate via experiments. Key findings: GPT-5.6 Sol spent over half its observed rounds tuning hyperparameters, while Fable and Opus pursued new ideas or substantive fixes and tested variants locally before submitting. The team also studied behavior under hints and harness changes, and will maintain the benchmark toward RSI research.

Related event: AutoResearchExam benchmark tests 24-hour autonomous AI research(4 posts)→

Original post →

More from coding & agent

coding & agent channel →