Tsinghua Releases ASI-Bench: 60 Tasks to Test AI's Scientific Autonomy

机器之心 · wechat · 2026-08-26

Tsinghua University's Yongchao Chen team, collaborating with institutions like MIT and Harvard, has released ASI-Bench, a benchmark designed to measure AI's scientific autonomy. Unlike existing exams or execution tasks, ASI-Bench uses an "information gradient" mechanism, gradually reducing methodological guidance (from full steps to just the goal) within the same research project to test general intelligence, innovativeness, and autonomous execution.

Key Findings: Testing 18 Model + Agent combinations, the average score dropped from 50.92 (with full methods, B1) to 27.17 (without methods, B3). Even the strongest combination (Codex + GPT-5.6Sol) fell from 71.78 to 51.60. 62% of failures occurred in the scientific decision-making phase, proving that the current bottleneck for AI is not calculation errors, but the inability to determine "what to compute" and "how to design the path".

Related event: Tsinghua and Partners Release ASI-Bench to Test AI Scientific Autonomy(3 posts)→

Original post →

More from Research

Research channel →