Mystery solved: Opus 5 vs 5.5 chart gap came from thinking levels, not regression
andrew_n_carr · x · 2026-09-23
Andrew Carr explains why the Opus 5.5 and Opus 5 Frontiercode charts seemed inconsistent: the two were measured under different conditions.
- The Opus 5.5 chart used adaptive thinking at Max, while the Opus 5 release chart used medium thinking.
- Opus 5 scores 48 at max but 53 at medium.
- Frontiercode penalizes out-of-scope/extra edits beyond what the model is told, so the thinking level directly shifts scores.
The numbers aren't a contradiction — just different test setups.
Related event: Opus 5's FrontierCode Score Drops from 53.4% to 48% in New Card(3 posts)→
More from Models
- NVIDIA releases Nemotron 3 Diarization model handling up to 8 overlapping speakers with 100M params — NVIDIAAI · 2026-09-23
- Dev slams Anthropic's Opus 5.5 safety checks for flagging basic code reviews — evilsocket · 2026-09-23
- Deep conversations with frontier models turn into incomprehensible AI-to-AI jargon, observer warns — erikphoel · 2026-09-23
- Andrew Carr: Opus 5.5 may be the first model that's a bit creative — andrew_n_carr · 2026-09-23
- Anthropic launches Life Sciences Verification Program to gate Opus 5.5 bio access — _sholtodouglas · 2026-09-23
- Rumor: OpenAI's GPT-6 'Astra Minor' is the new Sol, replacing Terra — daniel_mac8 · 2026-09-23