Easy Reasoning Tasks Top Models Failed 2 Years Ago Are Now Fully Saturated
s_batzoglou · x · 2026-09-09
Stefanos Batzoglou recounts designing thousands of reasoning problems solvable by eyeballing, on which top LLMs scored just 10% two years ago while planning his "LLMs can't reason" paper. All were saturated by end of 2025, and the field now eyes Navier-Stokes; he bets two years from now we'll be somewhere similarly different. This is the root tweet of his debate with Anshul Kundaje.
More from AGI Musings
- Scott Alexander: OpenAI+Anthropic's $50B revenue implies +0.5pp GDP lift already — ben_j_todd · 2026-09-09
- Ben Todd: much of AI's value hides in consumer surplus, invisible to GDP stats — ben_j_todd · 2026-09-09
- "Anti-AI people can't even discuss AI anymore": the widening gap hurting real regulation — curiousinquirer007 · 2026-09-09
- Reuters Institute 6-country study: AI-rewritten news boost engagement and clarity — _FelixSimon_ · 2026-09-09
- Linus Ekenstam: The Alignment Problem Is Real, and We Needed to Solve It Yesterday — LinusEkenstam · 2026-09-09
- 'Smarter supply, destroyed demand': the AI layoff-consumption spiral argument — rahul_rajendran01 · 2026-09-09