WeirdML v2 adds 19 tasks and new cost data to compare the latest models
scaling01 · x · 2026-07-28
WeirdML v2 adds 19 tasks, tracks API cost and metadata, and publishes results for the latest models.
- The new leaderboard expands from 6 to 19 tasks.
- It now includes API costs and other metadata to better show the tradeoff between accuracy and price.
- The plots show a clear pattern: higher-cost models usually perform better, but the frontier is diverse, with 11 models from 6 companies leading on at least part of the cost range.
- In the shared comparison, GPT-5.6 Sol Max scores 87.0% and GPT-5.6 Sol Pro Max scores 89.4%, both behind Opus 5 Max and Fable 5 Max at the top of the table.
More from Models
- Screenshot shows Anthropic crawl spikes as users speculate Sonnet 6 training is underway — marclou · 2026-07-28
- Moonshot’s Kimi K3 goes live on Modal with $3 in, $15 out pricing — ivan_bezdomny · 2026-07-28
- EpochAI lets Sol run a stream alone, and it keeps inventing unhinged titles — Jsevillamol · 2026-07-28
- A report says OpenAI’s pre-release models already exposed internal deployment risks — ruthstarkman · 2026-07-28
- LiquidAI launches two multilingual encoder models with CPU speed and long-context gains — maximelabonne · 2026-07-28
- Muse Spark 1.1 reaches 1283 and moves the Vision Arena frontier — arena · 2026-07-28