Independent eval of Opus 5.5 vs GPT-6 across 100 coding environments diverges from AAII

sandersted · x · 2026-09-24

A team tested Opus 5.5, GPT-6-Sol, and GPT-6-Luna across 100 open-ended coding/engineering environments, with results diverging significantly from AAII. GPT-6-Astra remains comfortably frontier. Opus 5.5 ranks #2-#7 on coding categories and is second-best at one-shot coding — their best correlate of fluid intelligence. Its lower pricing is welcomed as Anthropic's first reasonably priced model, but it underperformed in their custom harness, a pattern since Fable 5.1 where one-shot intelligence outpaces agentic results.

Original post →

More from Models

Models channel →