AutoResearchExam: a 24-hour benchmark finds AI research agents overfit, with Fable 5.1 edging Astra

AlexGDimakis · x · 2026-09-10

Alex Dimakis's team releases AutoResearchExam, a benchmark of open-ended ML and engineering tasks spanning seven research areas including training, data curation, safety and interpretability.

Related event: AutoResearchExam benchmark tests 24-hour autonomous AI research(4 posts)→

Original post →

More from coding & agent

coding & agent channel →