MIT and Harvard Study: LLMs Fall Short of Autonomous Scientific Discovery
davidmanheim · x · 2026-08-03
A joint paper by MIT and Harvard, Evaluating Large Language Models in Scientific Discovery, argues that current LLMs are nowhere near capable of autonomous scientific discovery.
Researchers developed a new evaluation framework called SDE, which takes LLMs out of traditional static multiple-choice tests and places them into real-world, open-ended research projects. The results show that despite weekly claims from tech labs about AI breakthroughs in biology, physics, or chemistry, there is close to no evidence that current AI can effectively assist in scientific discovery.
More from Research
- Top Information Retrieval Papers: Meta, YouTube, Tencent Explore RAG and Generative Recommendation — _reachsumit · 2026-08-03
- AI Curie Temperature Predictor Hits 6,000 Uses, Accelerating Material Screening — CatAstro_Piyush · 2026-08-03
- Deep Dive into Kimi K3: Architecture and Training of the 2.78T Model — imrancoder · 2026-08-03
- Classic Paper: Predicting Algorithm Runtime with Machine Learning by Hutter et al. (2014) — tomssilver · 2026-08-03
- Savage NeurIPS Rebuttal Title Roasts Reviewer for Not Reading the Paper — RylanSchaeffer · 2026-08-03
- AI-Assisted Pseudo-Proof Exploits Soundness Bug in Lean Kernel — AlexKontorovich · 2026-08-03