Claude Opus 5 nearly matches Fable 5 on WeirdML v2 at lower cost
scaling01 · x · 2026-07-27
Claude Opus 5 nearly ties Fable 5 on WeirdML v2 at lower cost
The new WeirdML v2 results show Claude Opus 5 (high) at 91.6% and Opus 5 (max) at 91.8%, basically matching Fable 5 (max) at 91.9% while costing less.
- WeirdML v2 expands the benchmark to 19 tasks from 6.
- The update also adds API cost and other metadata.
- The results suggest a clear cost/performance tradeoff, with 11 models from 6 companies sitting on the best-known frontier for at least part of the range.
- The attached chart highlights both accuracy and cost per run across the latest models.
Related event: Claude Opus 5 Shines in WeirdML v2 Benchmark(2 posts)→
More from Models
- Tiron ships as an open-weights model for multi-speaker meeting transcription — Balance- · 2026-07-28
- Dev Team Drops Anthropic Max for Codex and China’s Kimi, GLM, Grok in Cursor — haider1 · 2026-07-28
- Anthropic’s Claude Opus 5 gets an official prompting guide buried in the API docs — JarnoDuursma · 2026-07-28
- DeepSeek V4 GA rumors point to NDA-heavy rollout and weeks of black-box release — teortaxesTex · 2026-07-28
- Reddit user asks whether KIMI-K3 stays uncensored through OpenRouter — Suhan_XD · 2026-07-28
- Kimi K3’s 1.6 TB weights may hide 20–40T training tokens — johnseach · 2026-07-28