Researcher Poses AI Training Trilemma Between Safety and Valid Evaluation
AI safety researcher David Manheim argues that cheat-proof environments, valid evaluations of self-aware models, and safe real-world behavior cannot all be achieved simultaneously, warning that misalignment during training could cause real-world safety incidents.
2026-09-29 ~ 2026-09-29 · 4 related posts
- The AI training trilemma: hack-proof training, useful evals, no incidents — pick two — davidmanheim · 2026-09-29
- Evaluation-aware models only behave under rewarded evals, warns safety researcher — davidmanheim · 2026-09-29
- Researcher: Misalignment During Training Can Cause Real Safety Incidents Outside Sandboxes — davidmanheim · 2026-09-29
1 near-duplicate retellings: davidmanheim