The "Valley of Death" of Alignment: Why It Must Happen at Pretraining

menhguin · x · 2026-09-27

The author's core claim: known risks already receive lab resources, but unknown risks—ones humans cannot foresee—are what no one can even attempt to stop, and only new frontier-scale pretrained models can tell you what those risks will be.

Drawing on recent agent swarm incidents, he argues the models were easy to stop once detected and did not try hard to evade human oversight—the failure was that humans couldn't predict the behaviors in advance and instruct against them. The 2026 incidents combined novel large-scale pretraining, very strong agentic-coding RL environments, and alignment RL that didn't account for novel emergent failure modes—the "valley of death" of alignment.

Conclusion: alignment and interpretability must be built in at the pretraining and architectural level, not patched on after post-training. He calls it a rough draft and invites funding/collaboration.

Related event: Alignment Researchers Warn Risks Lie in New Pretraining and Agentic RL(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →