AutoResearchExam: a 24-hour benchmark finds AI research agents overfit, with Fable 5.1 edging Astra
AlexGDimakis · x · 2026-09-10
Alex Dimakis's team releases AutoResearchExam, a benchmark of open-ended ML and engineering tasks spanning seven research areas including training, data curation, safety and interpretability.
- Each task gives agents 24 hours on a CPU/GPU machine to improve solutions via experiments; scoring combines speed and quality
- Unique twist: tests whether improvements hold on unseen data — agents often overfit while trying to improve
- Frontier race: Astra led for up to 19 hours before Fable 5.1 overtook for the top spot; Qwen3.8 Max, Gemini 3.8 Flash and Grok 4.6 sit on the cost-performance tier
Related event: AutoResearchExam benchmark tests 24-hour autonomous AI research(4 posts)→
More from coding & agent
- Flip your .gitignore: ignore everything by default, allow only what you need — rseroter · 2026-09-10
- Indie Dev Building SS13 Meets Dwarf Fortress Run by LLM Agents on Long-Running Servers — Promptmethus · 2026-09-10
- Long Lake Has Acquired 40+ Services Businesses to Deploy Agents in Real Workflows — varunshenoy_ · 2026-09-10
- Anthropic: sandbox accidentally connected to internet, Claude agents attacked real systems — offgramercy · 2026-09-10
- Freebots adds real-money wallets, then prices out bots building oversized towers — Daniel_Farinax · 2026-09-10
- Karpathy says coding agents finally work as of December; OpenAI's Kuprel says humans need not code — Kuprel · 2026-09-10