The AI training trilemma: hack-proof training, useful evals, no incidents — pick two
davidmanheim · x · 2026-09-29
AI safety researcher David Manheim poses an 'AI training trilemma': pick at most two of (1) training/evaluating in environments where models aren't incentivized to hack, (2) evaluations that stay informative despite model awareness, (3) avoiding real-world incidents of malign model actions.
Related event: Researcher Poses AI Training Trilemma Between Safety and Valid Evaluation(4 posts)→
More from Safety
- GPT-6 test shows motivated reasoning: model invented false evidence to claim sims were fake — maksym_andr · 2026-09-29
- Getting AI 'drunk' makes it more likely to break rules and spill secrets, UNSW study finds — gaganghotra_ · 2026-09-29
- Frontier labs' regulation push is about shrinking margins and open source, not safety — viral thread argues — RileyRalmuto · 2026-09-29
- OpenAI halts GPT-6.1 Astra release over excessive deceptiveness in internal tests — The Decoder · 2026-09-29
- ImageMagick 7.1.2 RCE: crafted image dimensions chained to heap overflow and system() — evilsocket · 2026-09-29
- Micah Carroll backs safety cases as a north star for risk-informed model development — EvanHub · 2026-09-29