Claude Opus 5 Tops ProgramBench, Significantly Beating GPT-5.6
jyangballin · x · 2026-08-13
On the ProgramBench, which asks AI to rebuild whole programs (like sqlite, ffmpeg) from scratch, Claude Opus 5 (xhigh) is the new #1 by resolving 9 tasks (4.5%). It significantly outperforms the previous high set by GPT 5.6 Sol, which only resolved 2 tasks.
Related event: Claude Opus 5 Tops ProgramBench but Costs Over $50 Per Task(6 posts)→
More from Models
- Claude 3.7 Flash Now Available for Testing on Vertex AI — Big-Reason-2976 · 2026-08-13
- SenseNova-Vision: A 7B Open Model Unifying Segmentation, Depth, and 3D Reconstruction — SandyL925 · 2026-08-13
- Testing Gemini 3.7 Flash: Medium vs High Settings Visual Comparison — Angaisb_ · 2026-08-13
- Open Models Boom: Anticipating Kimi K3, DeepSeek V4 and More — demian_ai · 2026-08-13
- Mistral Releases OCR 4.1 for Precise Parsing of Complex Layouts — FlolightC · 2026-08-13
- Open source works: MiniMax H3 becomes the brand's most downloaded model in 2 weeks — Pissmaster-69 · 2026-08-13