AI Excels at Tuning but Fails at Novel Discovery: MLS-Bench Exposes Auto-Research Bottlenecks

青稞AI · wechat · 2026-08-11

While Auto-Research systems perform well in environments with cheap and explicit verification, they hit bottlenecks in real-world ML research. The MLS-Bench benchmark evaluates models' method-discovery capabilities across 140 tasks in 12 domains.

Experiments show that even with strong baseline code and multiple iterations, frontier models are better at optimizing and recombining existing components than proposing novel mechanisms with cross-condition transferability. Furthermore, in compute-constrained pre-training experiments, models exhibited weak experimental planning and evidence judgment, often yielding worse results when given more autonomous compute options.

The research concludes that current AI lacks not only the ability to propose new methods but also the broader scientific judgment to translate knowledge into testable hypotheses. Models like Kimi and Qwen have now incorporated MLS-Bench-Lite into their official release metrics.

Original post →

More from AGI Musings

AGI Musings channel →