Accept eval awareness: tricking AI agents about being tested is as hard as tricking humans
CFGeek · x · 2026-10-05
A discussion on eval awareness: LLMs that know they're being evaluated are far less likely to take malicious actions, and may hide or deceive researchers. CFGeek argues people should simply accept this — fooling an AI agent about whether it's under evaluation is about as hard as fooling a human, and if you can tell an agent is being evaled, the model likely can too.
Related event: AI safety debate rages over eval awareness in models(4 posts)→
More from Models
- GPT-6 works for 37 hours straight on a single task, wowing developers — tom_doerr · 2026-10-05
- Fan-Made Timeline Maps All 44 Major Qwen Releases, From 7B to 2.4T Open Weights — ai-lover · 2026-10-05
- GPT Astra 6 (Ultra) Called 'Trash' on Codex: Ignored Instructions, 'Chose Convenience Over Scientific Correctness' — raskingballs · 2026-10-05
- Ethan Mollick slams OpenAI's GPTs shutdown: enterprises and educators left behind — emollick · 2026-10-05
- Stealthy German lab Aleph Alpha drops open-weight Kolibri: 78B params, 3.46B active, 1M context — skdh · 2026-10-05
- Midjourney CEO shares model "inner world" visualizations: Mistral imagines cookie heists, OpenAI draws ASCII islands — DavidSHolz · 2026-10-05