Footnote reveals some Opus 5.5 eval scores were partly run by older Claude models
airesearch12 · x · 2026-09-23
A footnote in the Opus 5.5 launch eval table shows benchmarks ran with production safeguards on: when they intervened, cyber tasks were completed by Opus 4.8 and bio/frontier-LLM tasks by Opus 5 — meaning some Opus 5.5 scores partly reflect older models. Anthropic says this likely lowered the numbers, but the mixed-model scoring raises comparability questions.
More from Models
- Polymarket: GPT-6 claims up to 93% cheaper coding task costs than Claude Opus 5 — Polymarket · 2026-09-23
- OpenAI launches GPT-6 Sol and Luna, permanently cuts API prices 50% — aziz4ai · 2026-09-23
- Sam Altman announces GPT-6 Sol and Luna: major upgrades at half the price — sama · 2026-09-23
- Hugging Face tokenizers v1 runs in the browser via wasm — and it's fast — _lewtun · 2026-09-23
- GPT-6 Sol/Luna Share Astra's Cache Mechanics, Letting You Switch Effort Without Breaking Cache — brandon_galang · 2026-09-23
- GPT-6 Sol and Luna added to Artificial Analysis Intelligence Index v4.3 — ArtificialAnlys · 2026-09-23