MLS-Bench: AI Agents Can Optimize ML Experiments But Fail to Discover New Methods

机器之心 · wechat · 2026-08-12

Researchers from UC Berkeley, Tsinghua, and other institutions introduce MLS-Bench, a benchmark designed to evaluate whether AI agents can discover genuinely novel machine learning methods. Existing evaluations often fail to distinguish whether a score improvement comes from algorithmic innovation or mere engineering tweaks like hyperparameter tuning.

Evaluation Design and Core Mechanism

Key Findings: Strong Optimization, Weak Discovery

The benchmark has already been adopted in official evaluations for models like Kimi K3 and Qwen 3.8-Max.

Related event: Researchers Introduce MLS-Bench to Evaluate AI Research Capabilities(2 posts)→

Original post →

More from Models

Models channel →