Self-evolving Lean proof agents reach 45.1% on miniF2F with coevolving benchmarks

omarsar0 · x · 2026-07-22

A paper on self-modifying Lean proof agents argues that agents should co-evolve with their benchmarks instead of optimizing against a fixed test set.

The paper argues that verifier-grounded self-evolution can improve formal proof workflows while avoiding reward hacking.

Original post →

More from AGI Musings

AGI Musings channel →