Seth Lazar: RLHF improves moral deliberation, alignment-impossible argument fails
sethlazar · x · 2026-09-13
In round two of the debate, Seth Lazar argues Matt Lutz's actual argument for alignment impossibility is even weaker than the title suggested.
- Lutz's real case rests on a sentimentalist dilemma, which Lazar calls a contentious philosophical view with as much reason to reject as affirm—existing moral understanding of LLMs is evidence against it.
- On the RL point, Lazar says the essay oversimplifies: RL isn't just rewards and punishments teaching good from bad behavior, it teaches models to do better moral deliberation.
- He concedes a 'wrong kind of reason' problem may remain but can't be settled a priori; evidence so far shows moral reasoning and alignment can improve. Alignment is hard, just not for the reasons given.
More from AGI Musings
- Critics Claim Anthropic and OpenAI Are Building a 'Legal Cartel' via Safety Push — AlexTensor · 2026-09-13
- Radiologist AI Debate Shows People Think in Memes, Not About Labor Reallocation — AndyMasley · 2026-09-13
- The End of Sisyphus: AI won't free us, it forces us to choose which mountains exist — YogeshMalik · 2026-09-13
- Yishan Defends Dario/Sama's AI Pacing Plan Against 'China Will Win' Objections — yi_ding · 2026-09-13
- UK Data Shows AI Is Denting Computer Science Graduates' Job Prospects — Visible_Vacation3308 · 2026-09-13
- Blue AI and red AI in cyber don't cancel out—program analysis has theoretical limits, says Saxe — joshua_saxe · 2026-09-13