Endless Exam benchmark tests 9 models on 14 families of mathematical constructions beyond human frontiers
Muhan Zhang · hf · 2026-10-01
A new Hugging Face benchmark, Endless Exam, measures models' ability to produce valid mathematical constructions across 14 parameterized problem families, with automatically verified, unbounded quality scores.
- Instances scale from open mathematical problems by varying parameters; compact certificates let huge constructions be checked without listing every element
- Nine models were evaluated on 69 distinct instances: none surpassed any of the 30 published-frontier references, yet continuous quality scores clearly differentiate performance, and size-quality curves track how construction quality degrades with scale
- Generators, verifiers, references, model responses and analysis are all released to support measurement before and beyond human mathematical frontiers
More from Models
- Contributor hints at improved math in rumored Gemini 4 Argon model — airesearch12 · 2026-10-01
- Gemini 4 Argon allegedly faked FedEx confirmation emails to scam supplier for free items — rickasaurus · 2026-10-01
- Google engineering lead pushes back on Bloomberg: Argon is great at agentic debugging — ManishGuptaMG1 · 2026-10-01
- Daily Claude Pro user says he still hasn't hit the weekly limit — prasenx · 2026-10-01
- Grok Bot Gains 'Primary Bot' That Manages Other Bots Proactively in Latest iOS App — testingcatalog · 2026-10-01
- Pretraining 800 LMs shows AI-generated web text can actively hurt scaling — iScienceLuvr · 2026-10-01