Claude Opus 5 Shines in WeirdML v2 Benchmark
The upgraded WeirdML v2 benchmark now includes 19 tasks and API cost metadata. In this latest test, Claude Opus 5 scored 86.3%, nearly matching the performance of the Fable 5 model despite outputting over 7,000 tokens.
2026-07-26 ~ 2026-07-27 · 2 related posts
- Claude Opus 5 scores 86.3% on WeirdML v2 and still averages 7,000-plus tokens — xeophon · 2026-07-26
- Claude Opus 5 nearly matches Fable 5 on WeirdML v2 at lower cost — scaling01 · 2026-07-27