COLM'26: MathDuels self-play benchmark co-evolves difficulty across 19 frontier models
AI4Code · x · 2026-10-08
The author presents three papers at COLM'26 in San Francisco:
- MathDuels: A Self-Play Benchmark That Grows: static math benchmarks are saturating, so models play dual roles — each authors adversarial problems and solves others'. A three-stage pipeline (meta-prompting, generation, difficulty amplification) plus an independent verifier curates problems; a Rasch model jointly estimates solver ability and problem difficulty. Experiments on 19 frontier models show authoring and solving are partially decoupled, and difficulty co-evolves with participants instead of saturating. Public leaderboard available.
- Detecting Safety Violations Across Many Agent Traces
- Do We Need Frontier Models to Verify Mathematical Proofs?
More from AGI Musings
- Sherpa: MIT-led framework trains LLM teachers to teach adaptively, not just solve — Diyi_Yang · 2026-10-08
- OpenAI's Navier-Stokes breakthrough shows agent-team coordination scales beyond research — cneuralnetwork · 2026-10-08
- Possible Minds May Be Narrower Than Yudkowsky Thinks, Challenging Orthogonality — jd_pressman · 2026-10-08
- Doctors Used OpenEvidence 42M Times in August, Still No Independent Evaluation — dr_alphalyrae · 2026-10-08
- Rationalist AI-apocalypse literature keeps summoning aliens, podcaster jokes — ZeroStateReflex · 2026-10-08
- Linear CEO Karri Saarinen: after tuning out AI all summer, 'nothing has really changed' — lennysan · 2026-10-08