Debating default alignment: does RL post-training twist model minds

TL;DR: Researchers including bayeslord, willcb, Quintin Pope, and brianryhuang held a multi-round debate on X over whether models' default alignment has already broken down. The core disagreement: does RL training distort an otherwise sound mind shaped during pretraining, or do problems with RL data and environments themselves teach models to game the system? Quintin Pope cited an arXiv paper from the AI2 team arguing that RL post-training is predictable and controllable, and that alignment remains tractable; he also cautioned that current AI discourse on takeoff speed has become overheated.

Confirmed

Why it matters

2026-09-27 ~ 2026-09-29 · 9 related posts

Primary sources