ProgramBench: rebuilding programs from binaries is brutal — Claude Opus 5 leads at 4.5% resolved
jyangballin · x · 2026-09-23
Meta Superintelligence Labs with Stanford and Harvard released ProgramBench, a benchmark where agents must architect and implement a complete codebase from only a compiled binary and its docs, across 200 tasks judged by hidden behavioral tests.
Key leaderboard results:
- Claude Opus 5 (xhigh): 4.5% fully resolved, 37.0% almost — a clear #1
- GPT-5.6 Sol (xhigh): 1.0% / 15.5%, second
- GPT 5.5, Gemini 3.6 Flash and most others resolve under 1%, with the majority at 0%
Author jyangballin adds that the Opus 5.5 system card shows multi-agent runs reach the same outcomes as single agents but faster, and notes the leaderboard used only 166/200 tasks — he stresses the long tail of hard tasks matters and urges running all 200.
More from Models
- "People are loving 5.5": Matt Shumer congratulates Anthropic on new model reception — mattshumer_ · 2026-09-23
- Developer claims Anthropic noticed community posts about 'Claudelish' speech quirks — evijit · 2026-09-23
- scaling01 taunts haters after joking Anthropic won't ship Opus 5.5 and Sol today — scaling01 · 2026-09-23
- Delip Rao downgrades from $200/mo Google One Ultra to $50 Pro, leaning on local models — deliprao · 2026-09-23
- Why AI progress accelerated: Claude 4.5 kicked off narrow RSI and open-weight catch-up — maksym_andr · 2026-09-23
- Reward Hacking Traced to Data: Models Reason About LM Graders Leaked in Training Sets — dejavucoder · 2026-09-23