GPT Tops ProgramBench, But Long-Horizon Coding Remains Challenging
parth007_96 · x · 2026-08-11
OpenAI's GPT model (GPT 5.6 Sol xhigh) has taken first place on the ProgramBench benchmark for long-horizon software engineering, perfectly rebuilding 2 out of 200 programs (including cmatrix and a Rust hexdump viewer).
Despite the win, the overall resolution rate remains extremely low. The primary metric (fully resolved) sits at 1%, and the secondary metric (almost resolved) peaks at 16.5%, highlighting significant room for improvement in complex, long-horizon coding tasks.
Related event: GPT 5.6 Sol Tops ProgramBench, Halving Costs but Showing Python Bias(6 posts)→
More from Models
- Context Compacting Violates ToS? Developers Complain About Anthropic's Terms — nptacek · 2026-08-11
- DeepSeek Harness v4 Released with New Whale Logo — teortaxesTex · 2026-08-11
- Frustrated by Endless 'Cheap Model Hits Opus Level' Evaluation Posts — xeophon · 2026-08-11
- DeepSeek Experiences Slower Responses During Peak Usage Hours — ricklamers · 2026-08-11
- Muse Glimmer Lags in Agentic Evals, but Leads in Tool Use and Hallucination Control — ArtificialAnlys · 2026-08-11
- OpenAI gives cyber defenders a less-restricted new model — lofty23_smart · 2026-08-11