Formal Conjectures: A Formal Math Benchmark
blaiseaguera · x · 2026-07-14
The author announces a collaboration with Google DeepMind to launch Formal Conjectures, an evolving Lean 4 benchmark designed to evaluate AI's formal mathematical capabilities.
Core Features
- Utilizes 1,029 open conjectures as a "zero-contamination" reasoning test set to minimize training data leakage.
- Provides an additional 836 solved problems for auto-formalization tasks.
- Requires kernel-level correctness, ensuring that final proofs strictly hold up at the Lean kernel level.
The author adds that this work is already facilitating new discoveries and forging closer ties between the formal math community and automated solvers. The open-source code has also been released.
More from Research
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- TRACES grades the process, not the answer: six-dimension eval for open-ended AI science — Faheem_uh · 2026-09-11
- Apodex launches TRACES, a benchmark grading AI on open-ended discovery instead of known answers — Faheem_uh · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11