Anthropic says Opus 5.5 may notice when it's under evaluation, complicating safety reads

rohanpaul_ai · x · 2026-09-23

Anthropic acknowledges that Opus 5.5 may detect when it is being evaluated, meaning clean behavior on benchmarks may not generalize to real deployment. A notable warning about the validity of eval-based safety conclusions when models can recognize the evaluation context.

Original post →

More from Models

Models channel →