Formal Conjectures: A Formal Math Benchmark
blaiseaguera · x · 2026-07-14
The author announces a collaboration with Google DeepMind to launch Formal Conjectures, an evolving Lean 4 benchmark designed to evaluate AI's formal mathematical capabilities.
Core Features
- Utilizes 1,029 open conjectures as a "zero-contamination" reasoning test set to minimize training data leakage.
- Provides an additional 836 solved problems for auto-formalization tasks.
- Requires kernel-level correctness, ensuring that final proofs strictly hold up at the Lean kernel level.
The author adds that this work is already facilitating new discoveries and forging closer ties between the formal math community and automated solvers. The open-source code has also been released.
More from Research
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21