New Benchmark Tests Agents on Formal Proofs

OfirPress · x · 2026-07-09

Balaji Rao and colleagues introduced s2n-bignum-bench, a benchmark evaluating agents' ability to write proofs in HOL Light for cryptographic assembly code from Amazon's s2n-bignum repository. Codex 5.3 xhigh currently scores just 6.3%, indicating substantial room for improvement on this benchmark.

Original post →

More from coding & agent

coding & agent channel →