MLS-Bench Reveals: Frontier LLMs Still Lack True Methodological Innovation

新智元 · wechat · 2026-08-12

A new benchmark, MLS-Bench, covering 12 domains and 140 real-world tasks, delivers a pessimistic verdict on the research capabilities of current LLMs. Even when given the full implementation of human SOTA methods and allowed repeated experiments, frontier models fail to demonstrate reliable methodological innovation, mostly resorting to tuning and recombining known modules.

Key Findings:

MLS-Bench has already been adopted into the official release tables for KimiK3 and Qwen3.8-Max, signaling that AI for Research is the next major frontier for the industry.

Related event: MLS-Bench Shows LLMs Lack True Research Innovation(3 posts)→

Original post →

More from Models

Models channel →