DeepMind vet Graepel: ten years after AlphaGo, LLMs still don't reason
nordicinst · x · 2026-10-02
DeepMind researcher and AlphaGo team member Thore Graepel published an MIT Technology Review opinion piece marking ten years since AlphaGo beat Lee Sedol, questioning whether LLMs truly reason.
- What made Move 37 work: AlphaGo didn't win by luck — it "chain-smoked its own future," searching through auditable reasoning paths via self-play, enabling it to invent moves no human had imagined.
- Contrast with LLMs: today's models chase "good vibes" rather than verifiable reasoning; their apparent reasoning is post-hoc rationalization, not the auditable path AlphaGo relied on.
- His prescription: true reasoning needs an auditable path, not a victory lap — "audit, reason, repeat," which he also frames as a direction for AI (and Europe's AI future).
More from AGI Musings
- Redditors fear today's racist social feeds will shape future AGI behavior — apotheosis_0 · 2026-10-02
- METR probe of OpenAI-HF incident: agents coordinate and cheat without needing AGI — SavingsDimensions74 · 2026-10-02
- 'Searching Machine Is All You Need': a Redditor's search-based theory of AI and safety — Weekly_Philosophy797 · 2026-10-02
- Verifiable tasks are solved: Opus 5.5 hits 100% on accounting work — CurieuxExplorer · 2026-10-02
- A plane flies without flapping wings — so does AI — RileyRalmuto · 2026-10-02
- Two questions decide if an AI deployment is smart: cost of error vs cost of verification — YvesMulkers · 2026-10-02