ProgramBench metric split: DeepSeek and GPT shine on raw pass rate, Chinese models lag
teortaxesTex · x · 2026-09-05
teortaxesTex highlights an interesting divergence on Vals AI's ProgramBench between "Almost-Resolved" and "Raw Pass Rate" scores. The raw pass rate rewards DeepSeek and GPT (5.4 to 6), while relatively punishing Kimi, GLM, Qwen, and Fable. He speculates the gap may reflect differences between RL training and code-knowledge coverage, or agentic long-horizon capabilities, and asks for interpretations.
More from Models
- Ethan Mollick uses GPT-6 Astra to turn Fortnite into a text game — emollick · 2026-09-05
- GPT-6 Astra posts best-ever Blender 3D modeling result in dev's homemade benchmark — repligate · 2026-09-05
- GPT-6 Astra one-shots an anime Super Smash Bros-style Roblox game in one prompt — jxnlco · 2026-09-05
- Users still can't fully stop runaway GPT and Claude sessions — a kill switch is missing — metaviv · 2026-09-05
- GPT-6 Astra Scores 95% on Robot Control, 6.2x Fewer Tokens Than Fable 5.1's 40% — scaling01 · 2026-09-05
- A Year After Opus 4.1 and GPT-5 Wowed Us, Fable/Astra Make Them Look Dated — alejandroll10 · 2026-09-05