GPT Tops ProgramBench, But Long-Horizon Coding Remains Challenging

parth007_96 · x · 2026-08-11

OpenAI's GPT model (GPT 5.6 Sol xhigh) has taken first place on the ProgramBench benchmark for long-horizon software engineering, perfectly rebuilding 2 out of 200 programs (including cmatrix and a Rust hexdump viewer).

Despite the win, the overall resolution rate remains extremely low. The primary metric (fully resolved) sits at 1%, and the secondary metric (almost resolved) peaks at 16.5%, highlighting significant room for improvement in complex, long-horizon coding tasks.

Related event: GPT 5.6 Sol Tops ProgramBench, Halving Costs but Showing Python Bias(6 posts)→

Original post →

More from Models

Models channel →