Claude Opus 5 scores 86.3% on WeirdML v2 and still averages 7,000-plus tokens
xeophon · x · 2026-07-26
WeirdML v2 has launched with 19 tasks, up from 6, plus API cost and metadata tracking. In the quoted results, Claude Opus 5 (no thinking) scores 86.3% on the benchmark, placing it between GPT-5.5 (xhigh) and another top model on the leaderboard.
The new results also show that the model still emits more than 7,000 output tokens on average in no-thinking mode, and it is slightly more expensive than GPT-5.6 Sol (high). The broader benchmark page highlights a cost/performance frontier across many models from multiple vendors, with a strong relationship between higher spend and better accuracy.
More from Models
- A blunt defense of open-weight models says blocking others from releasing them is “evil” — rdesh26 · 2026-07-26
- Kimi K3 is expected to go open-weight tomorrow, boosting the open-source camp — Hot_Example_4456 · 2026-07-26
- Claude usage screenshot shows Max-plan limits and $1,775.98 in credits consumed — letandrewcook · 2026-07-26
- Open-weight 4B models approach o3-level performance on Swedish medical exams — AccomplishedCat4770 · 2026-07-26
- ChatGPT’s math output looks like a serious paper after a two-hour prompt — airkatakana · 2026-07-26
- Baseten’s paper writes 247 fake facts into Qwen3 and still can’t make them stick — gerardsans · 2026-07-26