Claude Opus 5 tops ProgramBench with 41.5% accuracy
scaling01 · x · 2026-07-27
A shared ProgramBench screenshot shows Claude Opus 5 taking the top spot on the coding benchmark, with 41.5% accuracy on the “Almost Resolved” task.
- Claude Fable 5 ranks second at 33.0%.
- GPT-5.6 Sol is third at 23.0%.
- The chart also lists Claude Opus 4.8, Claude Sonnet 5, GPT 5.5, and several other models, making it a snapshot of current coding-model performance rather than a product announcement.
More from Models
- Tiron ships as an open-weights model for multi-speaker meeting transcription — Balance- · 2026-07-28
- Dev Team Drops Anthropic Max for Codex and China’s Kimi, GLM, Grok in Cursor — haider1 · 2026-07-28
- Anthropic’s Claude Opus 5 gets an official prompting guide buried in the API docs — JarnoDuursma · 2026-07-28
- DeepSeek V4 GA rumors point to NDA-heavy rollout and weeks of black-box release — teortaxesTex · 2026-07-28
- Reddit user asks whether KIMI-K3 stays uncensored through OpenRouter — Suhan_XD · 2026-07-28
- Kimi K3’s 1.6 TB weights may hide 20–40T training tokens — johnseach · 2026-07-28