Humanize + GPT-5.5 solves 670/672 Lean-verified proofs, tops PutnamBench at 99.7%

songhan_mit · x · 2026-09-06

Ligeng Zhu's team claims their Humanize agent flow driving GPT-5.5 solved 670 of 672 formal statements on PutnamBench (99.7% pass rate) in about 5 hours, with every proof Lean-verified, taking #1 on the official leaderboard. Their key argument: as token quality converges across models, better agent orchestration is the lever for scaling agents. Note the result is self-reported and not yet independently reproduced.

Original post →

More from coding & agent

coding & agent channel →