Gary Marcus: Pure LLMs Aren't Good at Math, Credit Goes to Hybrid Systems

Gary Marcus recently posted multiple times, clarifying that the claim "LLMs are good at math" is misleading. He emphasized that LLMs cannot independently and reliably perform complex mathematical reasoning, and warned against overreacting without understanding the underlying technical details.

Confirmed

Marcus clearly distinguished between "pure LLMs" and "hybrid systems." He pointed out that modern mathematical AIs perform well on hard math tasks because they are actually hybrid systems combining neuro-symbolic tools. Merely scaling up model size cannot solve the inherent cascading hallucination problems in complex mathematical reasoning. Therefore, attributing the success of hybrid systems directly to pure LLMs is misleading. Furthermore, he used a car analogy to illustrate his point: just as knowing an engine's displacement isn't enough to judge a car's overall performance (requiring consideration of the transmission, suspension, etc.), decent performance on certain arithmetic tasks does not mean LLMs possess comprehensive reasoning abilities or genuine intelligence.

Why it matters

This perspective clarifies common misconceptions about the boundaries of LLM capabilities in the current AI industry. Amidst widespread hype over LLMs, explicitly pointing out the shortcomings of pure language models in rigorous logical reasoning helps the public and developers evaluate AI technology more objectively, highlighting the critical value of hybrid architectures, such as neuro-symbolic approaches, in achieving reliable reasoning.

2026-07-21 ~ 2026-07-23 · 5 related posts

Primary sources