ChatGPT and Claude Fail to Solve Math Proof Independently

2prime_PKU · x · 2026-07-07

Even when provided with the correct Harrison-Reiman subclass and key papers, ChatGPT 5.5 Pro and Claude Opus 4.8 failed completely to solve this mathematical proof independently in fresh conversations. This is the core conclusion of a Peking University research series: current cutting-edge frontier models still have a fundamental gap in hardcore mathematical reasoning, making human-AI collaboration a more realistic path than independent AI operation.

Related event: Peking University Team Tackles 35-Year-Old Queueing Theory Conjecture via Human-AI Collaboration(7 posts)→

Original post →

More from Models

Models channel →