ICML Paper: MATH-Perturb Reveals LLMs Rely on Pattern Matching in Math Reasoning
MengdiWang10 · x · 2026-08-25
Researchers from Princeton University and Google have released the MATH-Perturb benchmark and an accompanying paper (ICML 2025) to test the true mathematical reasoning capabilities of LLMs.
The study finds that when familiar math benchmark problems undergo small but fundamental structural changes—termed "Hard Perturbations"—model performance drops significantly. LLMs often continue down familiar solution paths (modes) rather than adapting to the new structure. This suggests that LLM performance on math tasks relies heavily on pattern matching and shortcuts rather than genuine generalizable reasoning. The benchmark consists of 279 perturbed problems derived from the hardest level-5 problems in the MATH dataset.
More from Research
- NVIDIA Releases ARDY: Real-Time Interactive Human Motion Generation Model — rsasaki0109 · 2026-08-27
- Menlo Achieves Zero-Shot Sim2real for Asimov Biped, Citing Unitree Early-2025 Parity — chris_j_paxton · 2026-08-27
- OpenWiki 0.4.2 Flaw Found: Adversarial Proof Reveals History Leak — tallmetommy · 2026-08-27
- Meitu MT Lab Presents CFT for Stable Portrait Relighting at ECCV 2026 — jiqizhixin · 2026-08-27
- Visual General Intelligence White Paper: Vision as a Pathway to AGI — zhenjun_zhao · 2026-08-27
- Expert Warns Against Self-Deception in ML Evaluation — PolarBearby · 2026-08-27