ProgramBench: rebuilding programs from binaries is brutal — Claude Opus 5 leads at 4.5% resolved

jyangballin · x · 2026-09-23

Meta Superintelligence Labs with Stanford and Harvard released ProgramBench, a benchmark where agents must architect and implement a complete codebase from only a compiled binary and its docs, across 200 tasks judged by hidden behavioral tests.

Key leaderboard results:

Author jyangballin adds that the Opus 5.5 system card shows multi-agent runs reach the same outcomes as single agents but faster, and notes the leaderboard used only 166/200 tasks — he stresses the long tail of hard tasks matters and urges running all 200.

Original post →

More from Models

Models channel →