Alignment debate: should we just accept eval awareness in AI safety testing?

RothRottweiler · x · 2026-10-05

An alignment-safety exchange: @CFGeek argues researchers should accept eval awareness — tricking AI agents into not knowing they're being tested is roughly as hard as tricking humans, so assume the agent knows whenever you do. @RothRottweiler pushes back: if you assume eval awareness, how can you run alignment evals at all? A real methodological tension for alignment research.

Related event: AI safety debate rages over eval awareness in models(4 posts)→

Original post →

More from Safety

Safety channel →