Tsinghua, MIT, Harvard launch ASI-Bench to measure AI's scientific autonomy
jiqizhixin · x · 2026-09-04
Tsinghua University, together with MIT, Harvard, CMU, USTC, Microsoft Research, and 10+ other institutions, has released ASI-Bench, the first benchmark designed to measure AI systems' scientific autonomy rather than recall of known knowledge. It evaluates whether AI can independently choose research questions, pick methods, decide what to observe, and self-correct when early steps fail—the combination of novelty and execution that defines real scientific discovery. The benchmark is open for community contributions.
More from Research
- Kangwook Lee: agent harness improvement can happen inside rollouts once agents are capable enough — Kangwook_Lee · 2026-09-04
- The Bayes Bandit: A Mathematical Take on Curiosity in Reinforcement Learning — CatAstro_Piyush · 2026-09-04
- Andriy Burkov Releases The Dark Hundred-Page Language Models Book on Kindle — burkov · 2026-09-04
- Alibaba-NLP CORE: boosting compositional reasoning in MLLM embeddings via reranker distillation — Alibaba-NLP · 2026-09-04
- On-policy distillation improves for hundreds of steps from a single query — it's algorithm-starved, not data-starved — Thinking-Space · 2026-09-04
- WorldReward: a vision-language reward model for evaluating camera-conditioned world models — Yibin Wang · 2026-09-04