Vero: First Benchmark for Repository-Scale Formal Verification by AI Agents
dawnsongtweets · x · 2026-08-23
Dawn Song's team introduces Vero, the first benchmark for joint implementation and proof synthesis at the repository level to evaluate AI's ability to build formally verified software. It includes 43 multi-module Lean 4 instances with 743 APIs and 2,705 specifications. Evaluations reveal that even the strongest frontier agent (GPT-5.5 at high reasoning effort) fully verified only 27 of 43 repositories within 90 minutes. The study identifies the construction of reusable lemma libraries for cross-module invariants as a major capability gap.
More from coding & agent
- Team of agents solves engineering problems end-to-end, manufacturing real objects — ProfBuehlerMIT · 2026-08-23
- Bezalel gives your AI agent memory, email, a computer and money behind one MCP URL — Rasmic · 2026-08-23
- Multi-Model Agent Workflow: GPT Writes, Claude Reviews, Auto-Generates PR — Saboo_Shubham_ · 2026-08-23
- Open Source Multi-Host Management for AI Jobs with Agent Orchestration — ii_social · 2026-08-23
- Agensis update auto-adds tasks from long-running requests — jasonkneen · 2026-08-23
- Open Source Windows MCP Connector Adds Sub-Agents and Computer Control to ChatGPT — Present-Boat-2053 · 2026-08-23