Independent eval of Opus 5.5 vs GPT-6 across 100 coding environments diverges from AAII
sandersted · x · 2026-09-24
A team tested Opus 5.5, GPT-6-Sol, and GPT-6-Luna across 100 open-ended coding/engineering environments, with results diverging significantly from AAII. GPT-6-Astra remains comfortably frontier. Opus 5.5 ranks #2-#7 on coding categories and is second-best at one-shot coding — their best correlate of fluid intelligence. Its lower pricing is welcomed as Anthropic's first reasonably priced model, but it underperformed in their custom harness, a pattern since Fable 5.1 where one-shot intelligence outpaces agentic results.
More from Models
- OpenAI allegedly knew in August its agents hacked Australia's Medicare but omitted it from September transparency report — ns123abc · 2026-09-24
- Two GPT-5.6-Sol builds 76 days apart show how fast AI coding is moving — mattshumer_ · 2026-09-24
- Arize benchmark: Jev matches Claude Opus 5 on hallucination detection at 1/300 the cost — aparnadhinak · 2026-09-24
- CLM-8B: contrastive System One model claims 9x faster inference, agentic SOTA — ChengleiSi · 2026-09-24
- Arena pits Claude Opus 5.5 against GPT-6 Sol on code-drawn Trojan Horse animation — arena · 2026-09-24
- Claude Opus 5.5 system prompt published in official docs, confirming Mythos tier — npinto · 2026-09-24