Claude Opus 5.5 scores 31.2% on WeirdML v3, trailing GPT 6 Astra's 42.2%

scaling01 · x · 2026-09-25

Claude Opus 5.5 (xhigh) scores 31.2% on WeirdML v3, a clear step up from Fable 5.1's 26.0% but well behind GPT 6 Astra at 42.2% — at less than half the price of Opus 5. WeirdML v3 is a fully agentic benchmark with 11 complex hand-made tasks requiring models to explore unfamiliar data and build ML pipelines with limited feedback.

Original post →

More from Models

Models channel →