Claude Opus 5 Shines in WeirdML v2 Benchmark

The upgraded WeirdML v2 benchmark now includes 19 tasks and API cost metadata. In this latest test, Claude Opus 5 scored 86.3%, nearly matching the performance of the Fable 5 model despite outputting over 7,000 tokens.

2026-07-26 ~ 2026-07-27 · 2 related posts