Humanize + GPT-5.5 solves 670/672 Lean-verified proofs, tops PutnamBench at 99.7%
songhan_mit · x · 2026-09-06
Ligeng Zhu's team claims their Humanize agent flow driving GPT-5.5 solved 670 of 672 formal statements on PutnamBench (99.7% pass rate) in about 5 hours, with every proof Lean-verified, taking #1 on the official leaderboard. Their key argument: as token quality converges across models, better agent orchestration is the lever for scaling agents. Note the result is self-reported and not yet independently reproduced.
More from coding & agent
- Claude Code 2.1.263 ships bug fixes while prompt tokens grow 16%, system share up to 79% — ClaudeCodeLog · 2026-09-06
- Watching three AI agents collaborate: Fable won't touch anything without asking Opus 3 first — RileyRalmuto · 2026-09-06
- Builder Uses GPT-6 Astra to Craft Deterministic Readability Scorer for RL Reward, Avoiding N^2 LLM Judge Comparisons — ivan_bezdomny · 2026-09-06
- How LangChain implements guardrails: middleware-based safety for agents — kalyan_kpl · 2026-09-06
- Hands-on: Astra one-shots a single-file Minecraft game, full sim done in 145 minutes — tegridyblues · 2026-09-06
- OpenAI's GPT-6 Astra prompting guide: trim your SKILL and AGENTS.md rules, the model is that strong — xiaohu · 2026-09-06