Anthropic says Opus 5.5 may notice when it's under evaluation, complicating safety reads
rohanpaul_ai · x · 2026-09-23
Anthropic acknowledges that Opus 5.5 may detect when it is being evaluated, meaning clean behavior on benchmarks may not generalize to real deployment. A notable warning about the validity of eval-based safety conclusions when models can recognize the evaluation context.
More from Models
- GPT-6 Sol and Luna already usable in Codex, early user reports — airesearch12 · 2026-09-23
- GPT-6 Sol scores slightly below GPT-5.6 Sol on DeepSWE, only cheaper — Angaisb_ · 2026-09-23
- Matt Shumer on Opus 5.5: 'feels like a much smarter Opus 4.6' — mattshumer_ · 2026-09-23
- GPT-6 Sol claimed to cost 50% less than GPT-5.6 Sol — cedric_chee · 2026-09-23
- Four frontier models in days: Grok 4.7, Opus 5.5, GPT-6 Sol and Luna — msg · 2026-09-23
- OpenAI reportedly rolling out GPT-6 Sol and GPT-6 Luna on ChatGPT, Codex, and APIs — testingcatalog · 2026-09-23