Claude Opus 4.8 hits 16.5% on ProgramBench’s almost-resolved metric
jyangballin · x · 2026-07-25
The post highlights a ProgramBench result showing Claude Opus 4.8 (xhigh) reaching a new high of 16.5% on the "almost resolved" metric.
- The same family of models is shown climbing from Opus 4.6 to 4.7, 4.7-xhigh, and then 4.8-xhigh.
- ProgramBench is a code-rebuild benchmark where models work until they decide they are finished, so the metric is about how often they get extremely close to full resolution.
- The attached chart frames the result as "more capable, not more effort," suggesting better efficiency as well as better score.
Related event: Claude Opus 4.8 Hits Record 16.5% on ProgramBench(3 posts)→
More from Models
- Kimi K3’s architecture is public, and the draft diagram shows KDA plus AttenRes — AccBalanced · 2026-07-25
- A user says $200/month Codex is so productive they now pay OpenAI $1,200 a month — robleclerc · 2026-07-25
- A Qwen3.6-27B merge blends reasoning and coding into one 27B model — pbaylies · 2026-07-25
- Google still indexes an Amazon Bedrock doc that mentions Claude Opus 5 — Angaisb_ · 2026-07-25
- User says Laguna S2.1 is too slow for planning but unusually strong at complex debugging — Prudent-Objective852 · 2026-07-25
- Antirez says the “Chinese frontier models are mostly distillation” story is wrong — antirez · 2026-07-25