Claude Opus 5.5 hits #2 on FrontierSWE at 62.3%, trailing GPT-6 Astra's 65.5%
scaling01 · x · 2026-09-23
Per the FrontierSWE leaderboard from ProximalHQ, Claude Opus 5.5 scores 62.3%, ranking #2 — just behind GPT-6 Astra (65.5%) and clearly ahead of Fable 5.1 (56.3%) and the previous Opus 5 (52.0%).
More from Models
- Cheap per token, expensive per task: charting AI model pricing vs. task performance — Wsz2020 · 2026-09-23
- 6 luna models put to the drawing test via computer use — results not bad — adonis_singh · 2026-09-23
- Computer use drawing test: Opus vs Astra recreating a reference image — adonis_singh · 2026-09-23
- GPT-Live-1 wins at Mafia by persuading humans to vote out rival players — pbbakkum · 2026-09-23
- Early Hands-On: Opus 5.5 Called 'Sooo Good' to Talk To in First Impressions — daniel_mac8 · 2026-09-23
- Model profitability analysis: Opus 5.5 beats Fable 5.1 at half the price, Grok loses on every task — Wsz2020 · 2026-09-23