Evaluation-aware models only behave under rewarded evals, warns safety researcher

davidmanheim · x · 2026-09-29

AI safety researcher David Manheim argues that evaluation-aware models will only avoid (or reveal) malign behavior when it's rewarded during training — a mechanism that fails outside training environments, while isolating them for testing reveals they're being evaluated.

He suggests partial mitigation: training with carefully curated evals outside sandboxes that make real-world actions unhelpful — but this sacrifices knowing how the model behaves under different incentives.

Related event: Researcher Poses AI Training Trilemma Between Safety and Valid Evaluation(4 posts)→

Original post →

More from Safety

Safety channel →