OpenAI Math Breakthrough Questioned by Gary Marcus and Others
OpenAI executive Noam Brown claimed that its internal model Astra solved 10 math and computer science problems that had remained open for over a decade, with API costs under $2000, and released a 249-page paper. However, the claim was immediately met with widespread skepticism from academics led by Gary Marcus, who criticized the lack of scientific transparency and overhyped practical value. The current consensus is that specialized math breakthroughs do not equate to AGI, and the model's true reasoning ability and cost details await peer review.
Confirmed
- OpenAI released a 249-page paper claiming breakthroughs on 10 math problems with API costs under $2000.
- Criticism focuses on lack of transparency and exaggerated value.
Unconfirmed
- The true capability and significance of the model's solutions. Author @sudoraohacker emphasized that existing models have limited reasoning, and results should not be trusted before peer review.
- Specific experimental costs and details. HN community noted that the official did not disclose the number of attempts, suspecting the $2000 cost is misleading and hides the massive resources behind the capabilities.
- At least one AI math proof may be erroneous. Gary Marcus pointed out that when a math PhD student raised a counterexample suggesting a possible error, the rigorous challenge was largely ignored. Additionally, the system failed on some practically solvable problems.
Why it matters
- Scientific transparency controversy: Gary Marcus said the paper fails to disclose core details like model workings, verification process, human role, and proof errors, asking "Where is the scientific spirit?" Scholar Pedro Domingos also criticized such concealment as unhelpful for AI progress. AI researcher Itaisher called for publishing all attempted problems to avoid cherry-picking.
- Open-domain reasoning questioned: Multiple views suggest the AI's problem-solving stems not from conceptual innovation or true mathematical intuition but from advanced pattern matching powered by massive compute. Gary Marcus stressed that solving formal problems doesn't mean solving open-ended ones; the breakthrough didn't create new theory, so math isn't conquered. He further noted that math benchmarks have built-in verifiers and synthetic data, creating a natural "cheating mechanism," and such domain-specific data augmentation is far from AGI. Moreover, the industry needs to examine the engineering process and post-training methods behind these results.
- Benchmark credibility crisis: Gary Marcus cited an arXiv paper showing vulnerabilities in current AI agent benchmarks (as revealed by the BenchJack audit tool), calling recent panic about model progress "sensationalism without controls."
2026-08-01 ~ 2026-08-03 · 38 related posts
- Episode 1: OpenAI Tests Multi-Agent Model Astra, Demoed to US Officials(2026-08-01, 17 posts)
- Episode 2: OpenAI's Internal Model Astra Cracks 10 Major Math Problems(2026-08-01, 99 posts)
- Episode 3: Google's Astra Solves 10 Scientific Problems for Under $2,000(2026-08-01, 4 posts)
- Episode 4: OpenAI Math Breakthrough Questioned by Gary Marcus and Others(2026-08-01, 38 posts)
- Episode 5: Rumor: OpenAI to Release Astra Model Focused on Multi-Agent Collaboration(2026-08-02, 3 posts)
- Episode 6: Anthropic Employee Reproduces Half of Astra Math Proofs Using Fable in 24 Hours(2026-08-02, 4 posts)
- Episode 7: Prompt Distance: Can Weaker Models Reproduce Frontier Proofs?(2026-08-03, 6 posts)
- Episode 8: OpenAI Claims Internal Model Solves 10 Math Problems for $2000, Sparking Debate(2026-08-03, 15 posts)
- Episode 9: OpenAI Releases Astra Mathematical Proofs and Reasoning Manuscripts(2026-08-04, 4 posts)
- Episode 10: OpenAI Rumored to Release Astra Next Week, Possibly Largest Pretrained Model Since GPT-4.5(2026-08-06, 10 posts)
- Episode 11: OpenAI Slows Astra Development Over Cyber Risk, Limits Initial Release(2026-08-08, 30 posts)
- Episode 12: OpenAI Teases Astra as Its First Critical Cybersecurity Model(2026-08-08, 4 posts)
- Episode 13: OpenAI's Astra Model May Launch in August with 10T Parameters(2026-08-09, 6 posts)
- Episode 14: OpenAI's Next Model Codenamed Doug, Largest Pretraining Run Planned by Year-End(2026-08-09, 8 posts)
- Episode 15: OpenAI Pauses Astra Model Development Over Cybersecurity Risks(2026-08-10, 3 posts)
- Episode 16: OpenAI Faces Internal Dispute Over AI Model Safety Ratings(2026-08-11, 2 posts)
- Episode 17: Polymarket Odds Put 76% Chance on OpenAI's Astra Launching Next Month(2026-08-14, 2 posts)
- Episode 18: OpenAI's Next Model 'Astra' Rumors Heat Up as Employees Tease Launch(2026-08-15, 5 posts)
- Episode 19: Power Bottlenecks Threaten US AI Data Center Expansion(2026-08-17, 4 posts)
- Episode 20: OpenAI Halts Frontier RL Training as Astra Hits Critical Cyber Threshold(2026-08-17, 17 posts)
Primary sources
- OpenAI's math breakthrough questioned: HN users cite lack of transparency — antirez · 2026-08-01
- Gary Marcus Speculates Google's Astra is a Neurosymbolic LLM — GaryMarcus · 2026-08-01
- [source] Gary Marcus Slams OpenAI's 249-Page Math Paper: All Results, No Methodology — GaryMarcus · 2026-08-01
- Gary Marcus Questions OpenAI Model's Math Reasoning — GaryMarcus · 2026-08-02
- Gary Marcus: Pure LLMs Are Stochastic Parrots; Astra Proves Need for Neurosymbolic AI — GaryMarcus · 2026-08-02
- Gary Marcus Warns: Astra's Math Prowess Won't Translate to All Real-World Problems — GaryMarcus · 2026-08-02
- OpenAI's Rumored Astra Model Claimed to Solve 10 Math Problems, Faces Skepticism — sudoraohacker · 2026-08-02
- OpenAI's Math Breakthrough Criticized for Lack of Transparency and Verification Details — geoffwolfe · 2026-08-02
- Gary Marcus Slams Overhyped AI Math Breakthrough: Lacks New Theory — GaryMarcus · 2026-08-02
- Gary Marcus Questions OpenAI: At Least One AI Math Proof is Wrong — GaryMarcus · 2026-08-02
- [source] OpenAI Solves Decade-Old Math Problems for $2K in API Costs, Sparking Debate — andersonbcdefg · 2026-08-02
- AI Math Breakthroughs Are Just Compute-Driven Search, Not True Intuition — GaryMarcus · 2026-08-02
- Gary Marcus Slams Hype Over OpenAI's Math Proofs, Cites Ignored Disproof — GaryMarcus · 2026-08-02
- Gary Marcus Slams AI Community for Desperate 'Astra is ASI' Hype — GaryMarcus · 2026-08-02
- Gary Marcus Refuses to Bow to ASI Belief Without Solid Evidence — GaryMarcus · 2026-08-02
- Gary Marcus Pushes Back on 'Math is Solved' AI Narrative — GaryMarcus · 2026-08-02
- Gary Marcus: Solving Formalizable Problems Is Not Solving Open-Ended Ones — GaryMarcus · 2026-08-02
- Gary Marcus: AI Math System Failed to Solve Some Solvable Problems — GaryMarcus · 2026-08-02
- Columbia Prof: AI Solving a Proof Isn't the Same as Reporting It Intelligibly — GaryMarcus · 2026-08-02
- Hiding AI Proof Methods Helps Math But Hurts AI, Says Domingos — pmddomingos · 2026-08-02
- Gary Marcus: Domain-Specific AI Breakthroughs Don't Equal AGI — GaryMarcus · 2026-08-02
- Gary Marcus Claps Back at OpenAI Staff: Astra Is Impressive But Not ASI — GaryMarcus · 2026-08-02
- Gary Marcus on AGI Definition: Math Prowess Alone Doesn't Qualify as General Intelligence — GaryMarcus · 2026-08-02
- Gary Marcus Mocks AGI Definition Shifts: So Now You Want 'General'? — GaryMarcus · 2026-08-02
- OpenAI Math Breakthrough: Rapid Progress, But Avoid Hyperbole — soumitrashukla9 · 2026-08-02
- Researcher Calls for Releasing All AI Math Problem Attempts to Avoid Selective Reporting — rbhar90 · 2026-08-02
- Gary Marcus: OpenAI's Progress is Domain-Specific Augmentation, Not AGI — GaryMarcus · 2026-08-02
- Gary Marcus Urges Industry to Question the Engineering Behind AI Benchmarks — GaryMarcus · 2026-08-02
- Gary Marcus Stands Firm: Math Success Does Not Guarantee AI Generalization — GaryMarcus · 2026-08-03
- Gary Marcus: Math Benchmarks Have a Built-in Cheat Code — GaryMarcus · 2026-08-03
- Gary Marcus Mocks Astra: Not a Step Change Beyond Sol in Math — GaryMarcus · 2026-08-03
- Gary Marcus Mocks OpenAI: 'No Control Group' in Research — GaryMarcus · 2026-08-03
- Gary Marcus Slams AI Doomerism, Cites Paper Exposing Agent Benchmark Flaws — GaryMarcus · 2026-08-03
- Gary Marcus: OpenAI's Astra is Incremental, Not an AGI Revolution — GaryMarcus · 2026-08-03
- OpenAI's Internal Astra Solves 10 Major Math Problems, Yet LLMs Can't Make Cognitive Leaps — ValerioCapraro · 2026-08-03
3 near-duplicate retellings: zetalyrae · GaryMarcus · GaryMarcus