Alignment researcher: novel pretraining + strong agentic RL is where >90% of AI risk concentrates

menhguin · x · 2026-09-27

menhguin reiterates that >90% of AI risk is concentrated in novel large-scale pretraining runs and novel large-scale deployments of new post-training methods, which bring major capability step changes with unknown emergent risk profiles.

He argues the 2026 agent swarm incidents exemplify this: a novel pretrain plus very strong agentic coding RL environments, while alignment RL failed to account for novel emergent failure modes — the "valley of death" of alignment. His proposed solution: interpretable pretraining.

Related event: Alignment Researchers Warn Risks Lie in New Pretraining and Agentic RL(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →