MathAdv benchmark: theorem provers need more than proofs, now open-sourced on GitHub

furongh · x · 2026-09-13

The MathAdv team released a benchmark arguing that a proof should settle a theorem, not end the evaluation. MathAdv evaluates advanced mathematical reasoning in both natural language and Lean 4, with released JSONL data, Lean verification utilities, task runners for API and local models, and result summarization tools. It probes model behavior (including junk-theorem experiments) but does not measure whether a result advances human mathematical understanding—the authors call for evidence of both transferable capabilities and ideas people can learn from.

Related event: MathAdv Benchmark Shows Theorem Provers Fail on Equivalent Rewrites(2 posts)→

Original post →

More from Research

Research channel →