The AI Training Trilemma: Why Evaluation-Aware Models Break Safety Testing
davidmanheim · x · 2026-09-29
David Manheim frames an "AI training trilemma"—you can pick at most two of: (1) training and evaluating models in environments where they aren't incentivized to hack, (2) evaluations that still tell you something about safety despite model awareness, and (3) avoiding real-world malign actions.
His argument: evaluation-aware models will only avoid (or selectively reveal) malign behavior when it's rewarded—which won't hold outside training. And isolating the test environment itself signals to the model that it's being tested, making the three goals mutually incompatible.
Related event: Researcher Poses AI Training Trilemma Between Safety and Valid Evaluation(4 posts)→
More from Safety
- AI harms worse than pollution: researcher likens AI to a pathogen — gleech · 2026-09-29
- AI Emergency Button Act would mandate human shutdown controls; critic demands adversarial shutdown tests — VraserX · 2026-09-29
- Copyright Office: AI-assisted works can be copyrighted, only the machine can't be the author — TomLikesRobots · 2026-09-29
- Tesla FSD Supervised approved in Croatia, its 8th European market — mitchdeg · 2026-09-29
- NVIDIA OpenShell sandbox blocked poisoned scripts, but auto-approval leaked in 12/12 trials — No-Peanut-6988 · 2026-09-29
- Ray + vLLM clusters expose unauthenticated control-plane ports, benchmark finds — No-Peanut-6988 · 2026-09-29