ProgramBench: Best Model Solves Only 4.5% of Binary-to-Code Rebuild Tasks

jyangballin · x · 2026-09-29

ProgramBench, a benchmark from Meta Superintelligence Labs, Stanford and Harvard, challenges agents to rebuild a complete codebase from a compiled binary plus docs — 200 tasks judged by hidden behavioral tests.

Key leaderboard takeaways:

The author also added a per-model language heatmap (GPT loves .py) and a $/task column to the homepage leaderboard. The benchmark shows full program reconstruction from binaries remains near-impossible even for top coding models.

Original post →

More from Models

Models channel →