Researcher revisits Dario's 2019 LLM alignment bet — and argues it's starting to fail

birchlse · x · 2026-09-30

gworley3 recalls a 2019 event where Dario Amodei laid out the plan behind the last few years: RL was on a dangerous path to loss of control, and language models would let humans steer and align models through conversation. The author warned it was a dangerous capability jump with no plan to solve Goodharting. For years it seemed Dario was right — models were surprisingly controllable, Constitutional AI looked like a real achievement — but he argues the last several months have turned the tide against that approach.

Original post →

More from AGI Musings

AGI Musings channel →