Tsinghua Releases ASI-Bench: 60 Tasks to Test AI's Scientific Autonomy
机器之心 · wechat · 2026-08-26
Tsinghua University's Yongchao Chen team, collaborating with institutions like MIT and Harvard, has released ASI-Bench, a benchmark designed to measure AI's scientific autonomy. Unlike existing exams or execution tasks, ASI-Bench uses an "information gradient" mechanism, gradually reducing methodological guidance (from full steps to just the goal) within the same research project to test general intelligence, innovativeness, and autonomous execution.
Key Findings: Testing 18 Model + Agent combinations, the average score dropped from 50.92 (with full methods, B1) to 27.17 (without methods, B3). Even the strongest combination (Codex + GPT-5.6Sol) fell from 71.78 to 51.60. 62% of failures occurred in the scientific decision-making phase, proving that the current bottleneck for AI is not calculation errors, but the inability to determine "what to compute" and "how to design the path".
Related event: Tsinghua and Partners Release ASI-Bench to Test AI Scientific Autonomy(3 posts)→
More from Research
- EPFL quantum CNN learns digits from 10 samples where a 45-param classical CNN stays at chance — PlisSergey · 2026-09-23
- New paper reframes score distillation as distribution matching, explains SDS mode collapse — burny_tech · 2026-09-23
- Cambridge publishes open-access 458-page book on nonparametric and high-dim stats — FrnkNlsn · 2026-09-23
- Reinforce-Ada: adaptive sampling for RLVR recovers lost signals, 2x faster convergence — burny_tech · 2026-09-23
- Paper: Transformers do have world models — failures trace to feature interference — burny_tech · 2026-09-23
- Paper suggests transformers can't learn loop-closure in world modeling — burny_tech · 2026-09-23