ProgramBench: Best Model Solves Only 4.5% of Binary-to-Code Rebuild Tasks
jyangballin · x · 2026-09-29
ProgramBench, a benchmark from Meta Superintelligence Labs, Stanford and Harvard, challenges agents to rebuild a complete codebase from a compiled binary plus docs — 200 tasks judged by hidden behavioral tests.
Key leaderboard takeaways:
- Claude Opus 5 (xhigh) ranks #1 with 4.5% fully resolved and 37% almost-resolved (≥95% tests passing), but at a steep $50.53 average cost per task
- GPT-5.6 Sol (xhigh) is #2 at 1.0% / 15.5% for just $6.08; GPT 5.5, Gemini 3.6 Flash and others trail
- Nearly every other model — Claude Opus 4.x, GLM-5.2, GPT 5.4 mini, etc. — has a 0% full-solve rate
The author also added a per-model language heatmap (GPT loves .py) and a $/task column to the homepage leaderboard. The benchmark shows full program reconstruction from binaries remains near-impossible even for top coding models.
More from Models
- OpenAI staff oddly relaxed as Claude Opus 5.5 hype builds, hinting at a counterpunch — haider1 · 2026-09-29
- Yacine: SWE benchmarks are the only ones people care about — total CS victory — yacineMTB · 2026-09-29
- Early Test: Opus 5.5 'Really Really Good' at Generating AWS Architecture Diagrams — amaarora · 2026-09-29
- Leaked figures claim to show attention details of Opus 5.5, unverified — Sauers_ · 2026-09-29
- Together AI cuts Qwen3.8-Flash pricing 40% for the rest of the month, targeting high-volume coding assistants — togethercompute · 2026-09-29
- Opus 5.5 and Sonnet 5.5 reportedly outperform Sol and Astra ahead of OpenAI's 20+ Dev Day launches — hibzy7 · 2026-09-29