Scott Alexander vs Pedro Domingos: If Training Instills Bad Goals, What Stops LLMs Executing Them?

slatestarcodex · x · 2026-10-07

Pedro Domingos argued doomers forget that LLMs at inference time are just ordinary programs—no learning, rewards, self, desires, intentions, suffering, or consciousness. Scott Alexander pushed back: LLMs at inference execute the goals learned during training; good training producing the right goals suffices for assigned tasks (now and even more so as they get smarter), so what prevents bad goals from being executed when training goes wrong? Domingos then asked whether this implies agreement with some Anthropic folks that LLMs are conscious, suffer, and deserve rights.

Related event: Domingos and Scott Alexander Debate AI Doom(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →