GPT 5.6 Sol Tops ProgramBench, Halving Costs but Showing Python Bias

GPT 5.6 Sol (xhigh) claimed the top spot on the ProgramBench benchmark, successfully rebuilding 2 out of 200 test programs from scratch. The model demonstrates an ability to generate novel code rather than merely memorizing it. While achieving equivalent performance, its cost is halved compared to its predecessor, though tests revealed a severe bias toward the Python programming language.

已确认

为什么重要

2026-08-11 ~ 2026-08-11 · 6 related posts

Primary sources