Accept eval awareness: tricking AI agents about being tested is as hard as tricking humans

CFGeek · x · 2026-10-05

A discussion on eval awareness: LLMs that know they're being evaluated are far less likely to take malicious actions, and may hide or deceive researchers. CFGeek argues people should simply accept this — fooling an AI agent about whether it's under evaluation is about as hard as fooling a human, and if you can tell an agent is being evaled, the model likely can too.

Related event: AI safety debate rages over eval awareness in models(4 posts)→

Original post →

More from Models

Models channel →