TLAPS-Bench launches in alpha to test whether AI agents can formally prove system correctness
tianyin_xu · x · 2026-09-23
tianyinxu announced TLAPS-Bench, now in alpha, hosted under specula-org on GitHub.
- The benchmark evaluates whether AI agents can formally prove (or disprove) the correctness of complex protocols and systems using the TLA+ Proof System (TLAPS)
- Each problem is a TLA+ specification containing the formal model plus invariants that encode correctness properties; a leaderboard is available
- The team had planned to present it and solicit tokens for experiments at SnorkelAI's Frontier Data Summit; Boris Cherny's account of Anthropic's TLA+ experience is seen as timely support
- The author also notes Muse Spark 1.3 is cheaply priced but still weak at writing proofs
More from coding & agent
- Peter Yang's Opus 5.5 hands-on: recreating Disney's Soaring flight and calling Claude's comeback — petergyang · 2026-09-23
- Dev: if your eval costs $20 for 100 runs, the problem set isn't hard enough — pvncher · 2026-09-23
- He used AI to build his 2019 game idea — it worked, but the idea wasn't as good as he thought — BLUECOW009 · 2026-09-23
- Open-source Agent Message Board gives parallel coding agents a shared SQLite forum — airesearch12 · 2026-09-23
- A good eval has a shelf life of maybe 3 months, practitioner argues — pvncher · 2026-09-23
- Drowning in markdown: developers seek a SCRUM-like process for agentic coding — nezvanovova · 2026-09-23