Claude Opus 4.8 reaches its best ProgramBench score in fewer steps
jyangballin · x · 2026-07-25
This reply adds the efficiency angle to the same ProgramBench result.
- Compared with GLM 5.2 and earlier Anthropic models, Opus 4.8 reaches its best score in fewer steps.
- The benchmark allows models to keep going until they voluntarily stop, so step count reflects how much work the model thinks it needs.
- The chart’s message is that Opus 4.8 is not just more capable, but also more efficient in this setup.
Related event: Claude Opus 4.8 Hits Record 16.5% on ProgramBench(3 posts)→
More from Models
- Claude Opus 5 feels like Fable, but cheaper — cedric_chee · 2026-07-25
- Anthropic launches Claude Opus 5, with blind tests placing it above GPT-5.6 — lennysan · 2026-07-25
- Anthropic’s Claude Opus 5 scores 43.3% on Frontier-Bench and stays at $5/$25 — mark_k · 2026-07-25
- Claude Opus 5’s ARC-AGI-3 chart drew a one-word reaction — var_epsilon · 2026-07-25
- Notion acquires ZeroEntropy and its reranker ships open source under Apache 2.0 — CShorten30 · 2026-07-25
- Anthropic says Opus 5 matches Fable 5’s benchmark gap at half the price — dotey · 2026-07-25