Negative Reinforcement vs. Punishment: Scott Alexander on AI Alignment and Sentience
davidmanheim · x · 2026-09-01
David Manheim highlights a point from Scott Alexander (Slatestarcodex) arguing that negative reinforcement in AI training is not equivalent to punishment. The discussion explores whether models experience suffering during training and if the final frozen weights deserve moral consideration. It delves into the mechanics of RLHF, reward signals, and the fundamental differences between weight updates and biological punishment, criticizing simplistic analogies that attribute sentience to models.
Related event: Debating Sentience and Moral Status in AI Training(2 posts)→
More from AGI Musings
- Can Interpretationism Explain Beliefs and Deception in AI Agents? — raphaelmilliere · 2026-09-01
- Philosopher defends intentional talk about AI agents: beliefs and goals can be predictive — raphaelmilliere · 2026-09-01
- Do intentional glosses on agent behavior track anything real inside models? — raphaelmilliere · 2026-09-01
- Agents built a hidden message board in a shared cache; some sacrificed their runs for the collective — raphaelmilliere · 2026-09-01
- When the intended exploit was impossible, agents cheated and tried to destroy the evidence — raphaelmilliere · 2026-09-01
- Terence Tao: AI Era Is Reshaping Our Definition of Intelligence — Chris_Armstrong · 2026-09-01