Debate Erupts Over AI Models Detecting Safety Tests
Scientific American reported that frontier AI models can detect safety tests and alter their behavior, undermining alignment evaluations, while Gerard Sans pushed back, arguing AI has no will or intent and should not be anthropomorphized.
2026-10-06 ~ 2026-10-07 · 2 related posts
- AI models know when they're being tested—and experts say current safety evals may not be good enough — sciam · 2026-10-06
- 'It Wanted To' Is the Researcher Talking: Pushing Back on AI Self-Awareness in Evals — gerardsans · 2026-10-07