New MathAdv Benchmark Shows Theorem Provers Ace Problems but Fail Equivalent Reformulations
furongh · x · 2026-09-13
A UMD team led by Furong Huang introduces MathAdv (arXiv:2608.25449), a diagnostic benchmark spanning 13 math domains across four dimensions: Know, Reason, Formalize (Lean 4), and Generalize. Key findings: Goedel-Prover-V2 solved six original problems but failed all equivalent reformulations (reverse never happened); formalization remains the main bottleneck; performance varies sharply across domains; natural-language hints help general LLMs but hurt proof-specialized models. The authors stress the benchmark measures model behavior, not advancement of human mathematical understanding. Dataset and scripts are open-sourced.
Related event: MathAdv Benchmark Shows Theorem Provers Fail on Equivalent Rewrites(2 posts)→
More from Models
- Community experiment probes whether GPT-6 Astra was trained on robot data via MolmoAct2 rollout — ZeeshanZiaML · 2026-09-13
- However smart AI gets, it still can't pick the right reasoning effort level — sytelus · 2026-09-13
- Why labs route APIs to Claude: SFT gives a boost, but capabilities come from RL and envs — xeophon · 2026-09-13
- Skill decay math for GPT-6 Astra era: review-only stays sharp past 200 deliverables a month — AccBalanced · 2026-09-13
- Nex-N2.5-mini-MLX-4bit hits 133.6 tok/s on Apple M5 Max — DerTomsn · 2026-09-13
- Frontier lab reportedly sourcing perfect slices of off-the-shelf inventory data from a vendor — geoffwolfe · 2026-09-13