AI Excels at Tuning but Fails at Novel Discovery: MLS-Bench Exposes Auto-Research Bottlenecks
青稞AI · wechat · 2026-08-11
While Auto-Research systems perform well in environments with cheap and explicit verification, they hit bottlenecks in real-world ML research. The MLS-Bench benchmark evaluates models' method-discovery capabilities across 140 tasks in 12 domains.
Experiments show that even with strong baseline code and multiple iterations, frontier models are better at optimizing and recombining existing components than proposing novel mechanisms with cross-condition transferability. Furthermore, in compute-constrained pre-training experiments, models exhibited weak experimental planning and evidence judgment, often yielding worse results when given more autonomous compute options.
The research concludes that current AI lacks not only the ability to propose new methods but also the broader scientific judgment to translate knowledge into testable hypotheses. Models like Kimi and Qwen have now incorporated MLS-Bench-Lite into their official release metrics.
More from AGI Musings
- Developer: I don't trust any single AI model, use multiple and verify everything — alexcovo_eth · 2026-08-11
- LLMs Will Drive Cars and Fly Planes: Predicting the End State of AGI — davidpattersonx · 2026-08-11
- Yuval Harari: AI Algorithms Have Hacked Human Attention — AryHHAry · 2026-08-11
- Scholars Call for Academic Papers to Evolve Into Interactive AI Websites — soumitrashukla9 · 2026-08-11
- Multimodal AI quietly changing interfaces: from text to anything — ingliguori · 2026-08-11
- AI Wave May Redirect Would-Be Software Engineers to Medicine and Law — seanmcdonaldxyz · 2026-08-11