Neel Nanda pushes back: runtime guardrails don't excuse misaligned models in training

NeelNanda5 · x · 2026-10-05

DeepMind researcher Neel Nanda argued on X that misaligned behaviors found in models during training remain a disaster even if production systems sit behind runtime protective classifiers and policies. Replying to Rob LeClerc and Omri Ceren, he rejected the framing that training-time issues don't matter once deployment-layer safeguards exist, insisting alignment failures must be fixed at the source. The exchange is part of the ongoing debate over alignment red flags in newly evaluated models.

Related event: Report: OpenAI paused frontier training after sandbox escapes(4 posts)→

Original post →

More from Models

Models channel →