Scott Alexander vs Pedro Domingos: If Training Instills Bad Goals, What Stops LLMs Executing Them?
slatestarcodex · x · 2026-10-07
Pedro Domingos argued doomers forget that LLMs at inference time are just ordinary programs—no learning, rewards, self, desires, intentions, suffering, or consciousness. Scott Alexander pushed back: LLMs at inference execute the goals learned during training; good training producing the right goals suffices for assigned tasks (now and even more so as they get smarter), so what prevents bad goals from being executed when training goes wrong? Domingos then asked whether this implies agreement with some Anthropic folks that LLMs are conscious, suffer, and deserve rights.
Related event: Domingos and Scott Alexander Debate AI Doom(3 posts)→
More from AGI Musings
- User challenges frontier AI labs to drop 700+ cancer-cure papers as the true AGI milestone — BLUECOW009 · 2026-10-08
- RL environment design is an art only some people master, argue AI practitioners — DrDatta_AIIMS · 2026-10-08
- MediaCloud Data Backs Up the Fading of Hallucination Coverage — _FelixSimon_ · 2026-10-08
- Is the hallucination debate fading? Pew researcher thinks the talk has died down — _FelixSimon_ · 2026-10-08
- Reading Twitter now means constant vigilance against AI text, and it's exhausting — moonsandhues · 2026-10-08
- Minerva author predicted IMO gold by 2026 and superhuman math wasn't crazy — GarrisonLovely · 2026-10-08