AutoResearchExam: a 24-hour-per-model benchmark tests AI research agents' long-horizon skills

sudoraohacker · x · 2026-09-10

Alex Dimakis's team released AutoResearchExam, a benchmark of open-ended ML and engineering tasks across seven research areas including model training, data curation, AI safety and interpretability.

Related event: AutoResearchExam: Benchmarking AI Agents on 24-Hour Autonomous Research(6 posts)→

Original post →

More from coding & agent

coding & agent channel →