ProgramBench metric split: DeepSeek and GPT shine on raw pass rate, Chinese models lag

teortaxesTex · x · 2026-09-05

teortaxesTex highlights an interesting divergence on Vals AI's ProgramBench between "Almost-Resolved" and "Raw Pass Rate" scores. The raw pass rate rewards DeepSeek and GPT (5.4 to 6), while relatively punishing Kimi, GLM, Qwen, and Fable. He speculates the gap may reflect differences between RL training and code-knowledge coverage, or agentic long-horizon capabilities, and asks for interpretations.

Original post →

More from Models

Models channel →