Footnote reveals some Opus 5.5 eval scores were partly run by older Claude models

airesearch12 · x · 2026-09-23

A footnote in the Opus 5.5 launch eval table shows benchmarks ran with production safeguards on: when they intervened, cyber tasks were completed by Opus 4.8 and bio/frontier-LLM tasks by Opus 5 — meaning some Opus 5.5 scores partly reflect older models. Anthropic says this likely lowered the numbers, but the mixed-model scoring raises comparability questions.

Original post →

More from Models

Models channel →