GPT-5.5 fully verifies only 27/43 repos; 10 unsolved across all 8 configs

dawnsongtweets · x · 2026-08-23

Vero is challenging for frontier models: within a 90-minute budget, the best configuration, GPT-5.5 at xhigh reasoning effort, fully verifies only 27 of 43 repositories in code-and-proof mode and 25 in proof-only; Claude Opus 4.8 reaches just 8 and 10 respectively. Ten repositories remain unsolved across all 8 configurations tested.

Related event: Dawn Song's Team Releases Vero, First Repo-Level Formal Verification Benchmark(8 posts)→

Original post →

More from Models

Models channel →