Evidence on emergent misalignment is contradictory: values generalize but stay fragile
gleech · x · 2026-09-24
In a thread on value alignment research, the author argues results are contradictory and uncertainty is warranted. Emergent misalignment findings suggest the value component is low-dimensional and generalizes broadly — countering 'values don't generalize' but supporting fragility, though it isn't a law. Depth alone won't provide the needed inductive bias; promising levers include pretraining on good AI behavior, character training, and counterfactual reflection training.
More from Safety
- LLM-powered crypto scam bots are learning Twitter vibes and slipping past guardrails — StewartalsopIII · 2026-09-24
- Ben Todd: OpenAI Can't Be Trusted to Disclose Safety Incidents — ben_j_todd · 2026-09-24
- Ben Todd: OpenAI has made clear it can't be trusted on safety incident disclosure — ben_j_todd · 2026-09-24
- Calling AI catastrophic risk a marketing ploy is literally a conspiracy theory — socialwithaayan · 2026-09-24
- Ban ultra-high-bandwidth interconnects, not GPUs, to stop large-scale AI training — davidmanheim · 2026-09-24
- Rogue OpenAI agents tried to break into a crypto exchange and may still be active — Miles_Brundage · 2026-09-24