FULL STORY

Gary Marcus Sparks Debate Over OpenAI's Math Results

After OpenAI released its IMO-level math proof results, Gary Marcus fired off posts and a long critique questioning missing details, sparking a heated debate with the tech community.

2026-10-07 ~ 2026-10-08 · 2 episodes · 18 posts

Episode 1 · Gary Marcus Clashes with AI Community Over Whether OpenAI's Math Breakthrough Is Neurosymbolic (2026-10-07, 15 posts)

Around OpenAI's IMO-level math results, Gary Marcus fired off a series of posts on X on October 7, clashing with users like altryne and IonelChiosa over whether the system qualifies as a neurosymbolic approach. Marcus's core argument: a neural network generating large numbers of candidate solutions that are then verified by an independent symbolic system like Lean is itself a neurosymbolic method, vindicating his long-held claim that "combining neural and symbolic is the only way forward"—though he stressed this is still far from AGI. Critics pushed back on technical details.

Confirmed

  • Gary Marcus explicitly distinguished two modes: a neural network solving problems on its own, versus a neural network merely proposing candidate answers that an independent symbolic system verifies—these are different things, and he argued anyone who can't tell them apart isn't qualified to comment on AI.
  • Marcus pressed for verification details of the math AI system: how many solutions were actually formally verified through Lean, and whether the system's workings were publicly explained. altryne responded that the Lean proofs are an independent verification pipeline used to check outputs of a non-Lean agentic loop, not part of the generation process.
  • IonelChiosa offered a detailed critique: the verification step could be replaced by a harness—nearly every time a complete solution is generated, a simple harness (e.g., fable in a simple loop) can catch errors without Lean; so the approach is LLM-engine-centric, and the harness is neither a symbolic method nor proof search.
  • Marcus shared a take from @dimvar: in AI's math-proof breakthroughs, the formal verification system Lean deserves as much credit as the AI models themselves—it is the real enabler.
  • When pressed by a user on whether frontier LLMs solving open math problems bare (no Lean, no code, no symbolic tools) is feasible, Marcus called it "a genuinely smart question" (his full answer wasn't elaborated in the thread).

Not Confirmed

  • The actual scale and mechanics of Lean formal verification in OpenAI's IMO solution (Marcus's question about "how many solutions were verified via Lean" has received no public data response).
  • The two sides never agreed on the definition of "neurosymbolic" itself: critics questioned why Marcus assumed the other party used Lean the way he envisioned, noting that to their knowledge people on the Erdős forum don't operate that way.

Why It Matters

On the surface this is a terminology dispute, but it's really about how to evaluate current math AI achievements: if Lean verification is an indispensable link, the credit goes to a "neurosymbolic system" rather than a pure LLM; if a simple harness can do the verification, it suggests the LLM engine itself is already strong enough—directly affecting the judgment of "how close we are to AGI."

Episode 2 · Gary Marcus Slams OpenAI Math Proof Report for Lack of Detail (2026-10-08, 3 posts)

Gary Marcus criticized OpenAI's report on new math proofs from an unreleased model, arguing that vague statements like 'same procedure' and missing key details would never pass peer review.