VeriSoftBench: repo-scale Lean 4 verification benchmark, best LLM scores just 41%
xiye_nlp · x · 2026-10-08
- VeriSoftBench, presented at COLM, is a new benchmark for repository-scale Lean 4 formal verification.
- It contains 500 proof obligations drawn from 23 open-source formal-methods repos, preserving real repo context and cross-file dependencies.
- Two eval modes (curated deps vs full repo context); frontier LLMs still struggle, with the best model reaching only 41.0% (curated) / 34.8% (full repo).
More from Research
- MIT left a coding agent on an island for 30 hours with no goal — it started to play — neuroecology · 2026-10-08
- Sherpa: MIT-led framework trains LLM teachers to teach adaptively, not just solve — Diyi_Yang · 2026-10-08
- Meta, DeepMind and Isomorphic join DOE/NIH Virtual Biology Initiative to model the cell — paulnovosad · 2026-10-08
- Is formalized code as verifiable math legit, or does it sound like a scam? — ziv_ravid · 2026-10-08
- Public Sector's Highest-Value Move: Build Big Datasets, Let AI Infer the Rest — Afinetheorem · 2026-10-08
- 9B and 0.6B Embedding Models Share One Vector Space for Encode-Small-Search — tomaarsen · 2026-10-08