ICML Paper: MATH-Perturb Reveals LLMs Rely on Pattern Matching in Math Reasoning

MengdiWang10 · x · 2026-08-25

Researchers from Princeton University and Google have released the MATH-Perturb benchmark and an accompanying paper (ICML 2025) to test the true mathematical reasoning capabilities of LLMs.

The study finds that when familiar math benchmark problems undergo small but fundamental structural changes—termed "Hard Perturbations"—model performance drops significantly. LLMs often continue down familiar solution paths (modes) rather than adapting to the new structure. This suggests that LLM performance on math tasks relies heavily on pattern matching and shortcuts rather than genuine generalizable reasoning. The benchmark consists of 279 perturbed problems derived from the hardest level-5 problems in the MATH dataset.

Original post →

More from Research

Research channel →