Vero: First Benchmark for Repository-Scale Formal Verification by AI Agents

dawnsongtweets · x · 2026-08-23

Dawn Song's team introduces Vero, the first benchmark for joint implementation and proof synthesis at the repository level to evaluate AI's ability to build formally verified software. It includes 43 multi-module Lean 4 instances with 743 APIs and 2,705 specifications. Evaluations reveal that even the strongest frontier agent (GPT-5.5 at high reasoning effort) fully verified only 27 of 43 repositories within 90 minutes. The study identifies the construction of reusable lemma libraries for cross-module invariants as a major capability gap.

Related event: Dawn Song's Team Releases Vero, First Repo-Level Formal Verification Benchmark(8 posts)→

Original post →

More from coding & agent

coding & agent channel →