Claude Opus 5 peaks at high reasoning level on Vibe Code Bench, then gets pricier
thesaraharminta · x · 2026-07-25
- A Vals.ai chart compares Claude Opus 5 across five reasoning levels on Vibe Code Bench.
- Accuracy rises from low (76.7%) to medium (82.0%) to high (89.8%).
- But the two highest settings, xhigh (88.3%) and max (88.4%), are slightly worse than high while costing significantly more per task.
- The takeaway: high appears to be the best tradeoff between accuracy and cost.
More from Models
- Can any AI really watch 24 fps video with strong comprehension? — mattshumer_ · 2026-07-25
- Kimi weights could turn the debate into hardware economics versus V4 — teortaxesTex · 2026-07-25
- Anthropic details Fable 5 orchestration, advisor mode, and cache costs — brada · 2026-07-25
- PaddlePaddle’s HPD-Parsing trends on Hugging Face as a document parsing model — PaddlePaddle · 2026-07-25
- Claude Opus 5 hits 30.2% on ARC-AGI-3, far ahead of GPT-5.6 Sol — rbhar90 · 2026-07-25
- Gemini Pro tells a user to search the web, then says it can’t access the content — gaganghotra_ · 2026-07-25