Claude Opus 5.5 scores 31.2% on WeirdML v3, trailing GPT 6 Astra's 42.2%
scaling01 · x · 2026-09-25
Claude Opus 5.5 (xhigh) scores 31.2% on WeirdML v3, a clear step up from Fable 5.1's 26.0% but well behind GPT 6 Astra at 42.2% — at less than half the price of Opus 5. WeirdML v3 is a fully agentic benchmark with 11 complex hand-made tasks requiring models to explore unfamiliar data and build ML pipelines with limited feedback.
More from Models
- OpenCode sitemap leaks GLM-5.5 Flash, GLM-5.4, Kimi K4, DeepSeek V4.1 Pro — teortaxesTex · 2026-09-25
- Open-source AI weekly: Xiaomi MiMo-V2.6-Pro tops open leaderboard at 1/45th of Opus 5's cost — 0xsachi · 2026-09-25
- Paste two AI plans into each other's threads and both will always say the other plan is better — DesignMike2020 · 2026-09-25
- Same Prompt Test: Claude Opus 5.5 Crushes GPT at Motion Graphics Video Design — thisiskp_ · 2026-09-25
- Gemini 4 likely landing next month as Google ships coding models every 3-4 weeks — haider1 · 2026-09-25
- Unverified demo claims GLM-5.3-Flash hits 240 tok/s in a single stream — SIGKITTEN · 2026-09-25