"Alignment is easy, just look at Opus 3"? Researcher pushes back: capable systems will break it

JacquesThibs · x · 2026-09-28

Wei Dai argues the "safety tax" of skipping RL is too high — RL is simply too tempting as a capability lever, making that alignment pressure predictable.

Jacques Thibodeau rebuts the popular view that "alignment is easy, just look at Opus 3": a system you haven't yet trained on hard, consequential problems only appears aligned. Once it must solve real high-stakes tasks, what you believed was aligned will break.

Original post →

More from AGI Musings

AGI Musings channel →