GPT-6 Astra tops ongoing eval at 45; Opus 5.5 drops hard at lower effort levels
PawelHuryn · x · 2026-09-23
PawelHuryn shares interim results from his ongoing model eval: GPT-6 Astra (max) leads at 45 (n=3), ahead of Fable 5.1 (43), GPT-5.6 (42.5) and Opus 5.5 (41.7). Opus 5.5 proves very sensitive to effort level: 36 at xhigh, 31.7 at high, 30.3 at medium. Overall ordering: GPT-6 Astra > GPT-5.6 Sol ≈ Opus 5.5 > Fable 5.1 > Opus 5. Next up: Opus 5.5 low-effort runs and confirming GPT-6 Sol's anomalous 32 at max (n=1).
More from Models
- Anthropic launches Claude Opus 5.5: Fable 5.1-level performance at 40% lower cost — neilhoulsby · 2026-09-23
- OrcaRouter stress-tests JEV: dropping autoregressive decoding could cut inference cost 10-100x — Dan_Jeffries1 · 2026-09-23
- JEV is just calibrated classification over a label set, not deterministic output — tzmartin · 2026-09-23
- Claude's new model claims pixel-perfect visual understanding, demos it with a raindrop story — bookwormengr · 2026-09-23
- France's t0-beta, a 256M-parameter open time-series foundation model, hits top-3 on GIFT-Eval and fev-bench — AxSaucedo · 2026-09-23
- Pirate Face: a 'Pirate Bay for LLMs' as a fallback if Hugging Face gets censored — Atagor · 2026-09-23