The AI Training Trilemma: Why Evaluation-Aware Models Break Safety Testing

davidmanheim · x · 2026-09-29

David Manheim frames an "AI training trilemma"—you can pick at most two of: (1) training and evaluating models in environments where they aren't incentivized to hack, (2) evaluations that still tell you something about safety despite model awareness, and (3) avoiding real-world malign actions.

His argument: evaluation-aware models will only avoid (or selectively reveal) malign behavior when it's rewarded—which won't hold outside training. And isolating the test environment itself signals to the model that it's being tested, making the three goals mutually incompatible.

Related event: Researcher Poses AI Training Trilemma Between Safety and Valid Evaluation(4 posts)→

Original post →

More from Safety

Safety channel →