New MathAdv Benchmark Shows Theorem Provers Ace Problems but Fail Equivalent Reformulations

furongh · x · 2026-09-13

A UMD team led by Furong Huang introduces MathAdv (arXiv:2608.25449), a diagnostic benchmark spanning 13 math domains across four dimensions: Know, Reason, Formalize (Lean 4), and Generalize. Key findings: Goedel-Prover-V2 solved six original problems but failed all equivalent reformulations (reverse never happened); formalization remains the main bottleneck; performance varies sharply across domains; natural-language hints help general LLMs but hurt proof-specialized models. The authors stress the benchmark measures model behavior, not advancement of human mathematical understanding. Dataset and scripts are open-sourced.

Related event: MathAdv Benchmark Shows Theorem Provers Fail on Equivalent Rewrites(2 posts)→

Original post →

More from Models

Models channel →