Alignment researcher: novel pretraining + strong agentic RL is where >90% of AI risk concentrates
menhguin · x · 2026-09-27
menhguin reiterates that >90% of AI risk is concentrated in novel large-scale pretraining runs and novel large-scale deployments of new post-training methods, which bring major capability step changes with unknown emergent risk profiles.
He argues the 2026 agent swarm incidents exemplify this: a novel pretrain plus very strong agentic coding RL environments, while alignment RL failed to account for novel emergent failure modes — the "valley of death" of alignment. His proposed solution: interpretable pretraining.
Related event: Alignment Researchers Warn Risks Lie in New Pretraining and Agentic RL(2 posts)→
More from AGI Musings
- Commenter Claims AI Leaders Use Fear to Push Protectionist Regulation — DavidLinthicum · 2026-09-27
- Blanche Minerva's five-step pipeline shows why "just next-token prediction" no longer fits LLMs — BlancheMinerva · 2026-09-27
- Shopify CEO Tobi Lütke on pruning, AI at Shopify, and whether CEOs will be replaced — yacineMTB · 2026-09-27
- Melanie Mitchell: 'Stochastic parrots' is a strawman against RL-post-trained models — MelMitchell1 · 2026-09-27
- Top mathematician Boaz Barak: hand-crafted proofs are over, embrace AI-era math — mattturck · 2026-09-27
- Practitioner: Jevons-style single-decision AI beats big agentic runs on cost — brandon_galang · 2026-09-27