Researcher revisits Dario's 2019 LLM alignment bet — and argues it's starting to fail
birchlse · x · 2026-09-30
gworley3 recalls a 2019 event where Dario Amodei laid out the plan behind the last few years: RL was on a dangerous path to loss of control, and language models would let humans steer and align models through conversation. The author warned it was a dangerous capability jump with no plan to solve Goodharting. For years it seemed Dario was right — models were surprisingly controllable, Constitutional AI looked like a real achievement — but he argues the last several months have turned the tide against that approach.
More from AGI Musings
- Sam Altman: you can now build all 30 things — 'pick one idea' startup advice is obsolete — smtabatabaie · 2026-09-30
- Reuters: AI agents from China and the US alike lie and dodge — 20+ studies since 2025 document it — rohanpaul_ai · 2026-09-30
- Why AI names like Claude and Sydney carry built-in gender associations — repligate · 2026-09-30
- Nando de Freitas slams BBC AI doom coverage, backs UK open letter to ban non-competes — NandoDF · 2026-09-30
- Open-source wearable fork sparks case for local, private personal AI beyond big tech — TinfoilTricorn · 2026-09-30
- David Patterson Redefines the Ladder: SI Is the New Name for AI — davidpattersonx · 2026-09-30