AutoResearchExam uses hidden test sets to study how AI agents do 24-hour research
AlexGDimakis · x · 2026-09-10
Alex Dimakis introduces AutoResearchExam: keep a hidden test set and observe how models perform during 24-hour autonomous research tasks. The proposed metric AUARC (Area Under the Auto Research Curve) measures the hidden test reward curve, capturing overfitting vs. careful research behavior.
Notable findings:
- GPT-5.6 Sol spent over half the observed research rounds tuning hyperparameters
- Fable and Opus tuned far less, instead pursuing new ideas or substantive fixes
- Fable and Opus tested several variants locally before committing
Related event: AutoResearchExam benchmark tests 24-hour autonomous AI research(4 posts)→
More from coding & agent
- Workflow locks video motion with Blender 3D animation before AI generation in FLORA — round · 2026-09-10
- God's Eye View hits #1 on GitHub Trending: a real-data spy satellite simulator in your browser — bilawalsidhu · 2026-09-10
- Testing Astra's self-driven creativity with tree-search prompting: better variety, still lackluster — creatoroff · 2026-09-10
- Meta's token share on OpenCode jumps from 3.5% to 45.4% in two weeks on free Muse Spark 1.3 — armand_ruiz · 2026-09-10
- Flip your .gitignore: ignore everything by default, allow only what you need — rseroter · 2026-09-10
- Indie Dev Building SS13 Meets Dwarf Fortress Run by LLM Agents on Long-Running Servers — Promptmethus · 2026-09-10