New Moravec's Paradox: LLMs Ace Math but Fail Everyday Tasks
Several industry observers and developers recently highlighted an absurd capability mismatch in current large models: they can solve complex, open-ended math problems but frequently fail at everyday tasks like modifying invoices or following simple instructions, needlessly wasting millions of Tokens and hundreds of dollars. This contrast reveals a massive gap between the models' abstract reasoning abilities and their capacity to handle messy real-world rules.
Confirmed
- Capability Mismatch: @YuchenjUW and @shakoistsLog pointed out that while AI Agents excel at advanced math, they frequently err when executing tasks like modifying client invoices or following 3 simple instructions, even consuming massive amounts of resources.
- Derailment in Long Tasks: @NYCounihan reported that models lack focus when writing code continuously for 5 hours, easily drifting off-topic and generating useless "AI garbage" side tasks.
Why It Matters
- The New Moravec's Paradox: @ajratner argues this is essentially a modern iteration of Moravec's paradox, where math and programming—perceived as highly difficult for humans—are easy for AI, yet simple everyday tasks are incredibly hard for it. He cautions against linearly extrapolating AI breakthroughs in math and coding to its performance in other areas lacking verifiable data.
- Lack of Proactive Questioning: @Franc0Fernand0 analyzed that the core gap between AI and human interns isn't pure intelligence. When encountering illogical or contradictory situations, humans instinctively pause to ask questions, a capability currently missing in AI, which is a key reason for its frequent failures in real-world business operations.
2026-08-02 ~ 2026-08-03 · 8 related posts
Primary sources
- AI Agents Ace Complex Math But Struggle to Update Invoices — shakoistsLog · 2026-08-02
- AI Will Solve the Riemann Hypothesis Before Doing Our Taxes — rickasaurus · 2026-08-02
- [source] The Irony of AI: Solves Open Math Problems, Fails 3 Simple Instructions — Yuchenj_UW · 2026-08-03
- AI Solves Fields Medal Math But Fails Routine Debugging, Expert Jokes — mboehme_ · 2026-08-03
- [source] AI's New Moravec's Paradox: Math and Coding Soar While Other Domains Lag — ajratner · 2026-08-03
- [source] AI Solves Math Problems but Fails Simple Instructions? The Gap is About Asking Questions — Franc0Fernand0 · 2026-08-03
- AI Can Solve Math Equations But Wanders Off During 5-Hour Coding Tasks — NYCounihan · 2026-08-03
- New Moravec's paradox: LLMs solve math but struggle to build a $100 app — paraschopra · 2026-08-03