Claude Opus 5 scores 86.3% on WeirdML v2 and still averages 7,000-plus tokens
xeophon · x · 2026-07-26
WeirdML v2 has launched with 19 tasks, up from 6, plus API cost and metadata tracking. In the quoted results, Claude Opus 5 (no thinking) scores 86.3% on the benchmark, placing it between GPT-5.5 (xhigh) and another top model on the leaderboard.
The new results also show that the model still emits more than 7,000 output tokens on average in no-thinking mode, and it is slightly more expensive than GPT-5.6 Sol (high). The broader benchmark page highlights a cost/performance frontier across many models from multiple vendors, with a strong relationship between higher spend and better accuracy.
Related event: Claude Opus 5 Shines in WeirdML v2 Benchmark(2 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23