MLS-Bench Reveals: Frontier LLMs Still Lack True Methodological Innovation
新智元 · wechat · 2026-08-12
A new benchmark, MLS-Bench, covering 12 domains and 140 real-world tasks, delivers a pessimistic verdict on the research capabilities of current LLMs. Even when given the full implementation of human SOTA methods and allowed repeated experiments, frontier models fail to demonstrate reliable methodological innovation, mostly resorting to tuning and recombining known modules.
Key Findings:
- Lack of Fundamental Innovation: Models excel at engineering (e.g., reading code, assembling components) but cannot answer why a new mechanism is needed or what fundamental flaw it addresses.
- Limits of Scaling Search: Increasing inference compute and sampling only optimizes existing directions; it cannot create a new search space and often leads to local hill-climbing.
- Compute & Knowledge Aren't Cures: Giving models larger experiment budgets often degrades performance, and providing web search or theoretical background yields limited improvements.
MLS-Bench has already been adopted into the official release tables for KimiK3 and Qwen3.8-Max, signaling that AI for Research is the next major frontier for the industry.
Related event: MLS-Bench Shows LLMs Lack True Research Innovation(3 posts)→
More from Models
- AI Safety Guardrails Block Proof of Irreducibility for Stern Polynomials — chaumian · 2026-08-12
- MiniMax H3 Dominates: 4 out of 8 Trendy Models on HuggingFace — Xianbao_QIAN · 2026-08-12
- Xiaomi's Open Multilingual Translation Model Outperforms Proprietary Baselines — xiaomi-research · 2026-08-12
- DeepSeek Prefix Cache Hacks: Cut Agent Token Costs by 90% to $0.005/Task — BodybuilderLost328 · 2026-08-12
- Local MoE Benchmark: NVIDIA Lightning Outruns Qwen by 2.5x — parepeg · 2026-08-12
- Dev Calls for Official MiniMax H3 Turbo as Community Floods HF with Fine-tunes — cocktailpeanut · 2026-08-12