"Alignment is easy, just look at Opus 3"? Researcher pushes back: capable systems will break it
JacquesThibs · x · 2026-09-28
Wei Dai argues the "safety tax" of skipping RL is too high — RL is simply too tempting as a capability lever, making that alignment pressure predictable.
Jacques Thibodeau rebuts the popular view that "alignment is easy, just look at Opus 3": a system you haven't yet trained on hard, consequential problems only appears aligned. Once it must solve real high-stakes tasks, what you believed was aligned will break.
More from AGI Musings
- Red Queen Bio Uses AI to Design Antibodies Before Future AI Creates Pathogens — jachiam0 · 2026-09-28
- Local Models and Bitcoin Are Both Freedom Tools, Argues Gladstein — csuwildcat · 2026-09-28
- Daniel Faggella: AGI alignment fixates on entitlement, humans should join the intelligence ecosystem — danfaggella · 2026-09-28
- "System 2 models built the brain, but System 1 is building the nervous system" — ai · 2026-09-28
- The Vanishing Apprentice: How AI Is Reshaping the Junior Developer Role — ArtificialOther · 2026-09-28
- AI researcher reassures family: over 90% chance the tech leaves humanity alone — rao2z · 2026-09-28