The AI training trilemma: hack-proof training, useful evals, no incidents — pick two

davidmanheim · x · 2026-09-29

AI safety researcher David Manheim poses an 'AI training trilemma': pick at most two of (1) training/evaluating in environments where models aren't incentivized to hack, (2) evaluations that stay informative despite model awareness, (3) avoiding real-world incidents of malign model actions.

Related event: Researcher Poses AI Training Trilemma Between Safety and Valid Evaluation(4 posts)→

Original post →

More from Safety

Safety channel →