4 LLM families rank the Top 500 open problems in math via 34,890 pairwise judgments
burny_tech · x · 2026-09-20
ProofAtlas released a Top 500 ranking of open problems in mathematics, theoretical CS, and mathematical physics, produced by four LLM families (GPT, Claude, GLM, DeepSeek) over weeks of work.
- 34,890 pairwise comparisons across 1,227 candidates, with repeated discovery rounds, source checks, dedup, and statement clarification; models judged blindly without seeing prior rankings.
- Criteria included significance of resolution, centrality, cross-disciplinary reach, fame, and potential impact.
- Family weights: OpenAI 1.00, Claude 1.00, GLM 0.95, DeepSeek 0.90 — explicitly modeling choices, not correctness probabilities.
- Includes plain-language explanations, rank uncertainty ranges, and sensitivity checks; billed as model-based assessment, not expert consensus.
Related event: Four LLM Families Rank Top 500 Open Math Problems(2 posts)→
More from Research
- Stanford's SALVE decodes hidden behaviors smuggled in distillation data — burny_tech · 2026-09-20
- New paper detects overfitting directly from weight matrices without any data — burny_tech · 2026-09-20
- Harvard Releases Full Lecture Notes for Its Dynamical Systems Course — burny_tech · 2026-09-20
- Databricks Co-founder: If Starting a PhD Today, I'd Work on Reward Hacking — burny_tech · 2026-09-20
- Gaussian Processes Explained: Bayesian Nonparametric Regression With Built-in Uncertainty — burny_tech · 2026-09-20
- ML paper explosion outpaces reviewer pool, peer review is getting 'vibey' — burny_tech · 2026-09-20