Claude Opus 5 Tops ProgramBench, Doubling Fully Resolved Tests
jyangballin · x · 2026-08-13
Claude Opus 5 (xhigh) takes the #1 spot on the ProgramBench benchmark, which tasks AI with rebuilding whole programs like sqlite and ffmpeg from scratch.
Every metric hit a record high: fully resolved tests (100% pass rate) jumped from 2 to 9 out of 200; near-perfect scores (>95% pass rate) increased from 33 to 74; and the average test pass rate rose from 70.9% to 74.7%.
Related event: Claude Opus 5 tops ProgramBench, cost over $50(7 posts)→
More from Models
- Google Launches Gemini 3.7 Flash: Big Jumps in Coding & Agents at Half the Price — koraykv · 2026-08-14
- Google Demos 3-Agent Team Autonomously Training Robotics Models with Gemini 3.7 Flash — koraykv · 2026-08-14
- Google Gemma 4 12B Uses Encoder-Free Architecture for Multimodal Inputs — MaartenGr · 2026-08-14
- OpenAI Previews Ultrafast Mode: GPT-5.6 Sol Hits 14x Speeds — OpenAI · 2026-08-14
- Web Dev AI Leaderboard: Claude Takes First, Kimi Follows Closely — arena · 2026-08-14
- Hands-on: Gemini 3.7 Flash is Now Accessible via API — Angaisb_ · 2026-08-14