Negative Reinforcement vs. Punishment: Scott Alexander on AI Alignment and Sentience

davidmanheim · x · 2026-09-01

David Manheim highlights a point from Scott Alexander (Slatestarcodex) arguing that negative reinforcement in AI training is not equivalent to punishment. The discussion explores whether models experience suffering during training and if the final frozen weights deserve moral consideration. It delves into the mechanics of RLHF, reward signals, and the fundamental differences between weight updates and biological punishment, criticizing simplistic analogies that attribute sentience to models.

Related event: Debating Sentience and Moral Status in AI Training(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →