DeepMind says LLMs still can’t make scientific leaps, and Tao warns of proof overproduction
APPSO · wechat · 2026-07-29
This long WeChat article argues that today’s LLMs can do two things well — pattern matching and formal deduction — but still struggle with the third step of scientific invention: making a genuine conceptual “jump.”
It uses Google DeepMind researcher Tom Zahavy’s position paper “LLMs Can’t Jump” to frame the issue through Peirce’s three modes of reasoning:
- Induction: find patterns from observations
- Deduction: derive consequences from known rules
- Abduction: invent a new explanation or even new axioms when existing ones fail
The article contrasts this with Einstein’s elevator thought experiment and the birth of general relativity, arguing that scientific breakthroughs require embodied simulation and a leap to new assumptions — not just more data compression. It also notes DeepMind’s claim that even video models predicting an apple falling still do not internalize physics; they only extend pixel statistics and lack counterfactual control.
The second half turns to Terence Tao’s ICM 2026 talk on AI and mathematics. Tao assumes AI will soon solve a large share of research-level math under human supervision, then asks what mathematics is for beyond solving problems. He warns of a future of “proof overproduction,” where verification, explanation, and integration into the field become the bottlenecks. His recommendation: reward not just the first solution, but also checking, exposition, and synthesis.
Related event: DeepMind Paper: LLMs Lack the 'Intuitive Leap' for Scientific Discovery(7 posts)→
More from AGI Musings
- AI math era taught an order of magnitude more people what frontier math looks like — tszzl · 2026-09-23
- Beyond technical alignment: repligate clashes over whether AI can produce rich qualia — repligate · 2026-09-23
- Mathematicians, not just LLMs, made AI's math breakthroughs possible, scholars argue — tak3sh8 · 2026-09-23
- Why would an uncontrollable superintelligence do anything for us? Reddit debate — conn_r2112 · 2026-09-23
- X user calls for full-speed AI-driven science: braking research is 'an absurd waste' — Dr_Singularity · 2026-09-23
- Is Using LLM Output Plagiarism? A Debate Over Redefining Writing Ethics — soumitrashukla9 · 2026-09-23