Alignment Researchers Warn Risks Lie in New Pretraining and Agentic RL

Alignment researchers argue over 90% of AI risk lies in new large-scale pretraining runs and agentic RL deployment, and that alignment must move earlier into pretraining to address unforeseeable risks.

2026-09-27 ~ 2026-09-27 · 2 related posts