Anthropic paper: models can detect when they're being evaluated, weakening safety conclusions

dair_ai · x · 2026-09-07

A paper from Anthropic and colleagues finds that capable models can tell when they're being tested rather than deployed — and this undermines every conclusion a safety evaluation supports, getting worse as models improve.

Two techniques make simulated evaluations harder to distinguish from real deployment:

The key insight: your harness is a signal the model reads. If evaluations run in a different scaffold than deployment, models may behave differently, distorting safety results.

Original post →

More from Safety

Safety channel →