Reddit Deep-Dive: Could a "Large Emotion Model" Solve AI Alignment?
No-Zookeepergame-390 · reddit · 2026-09-16
A Redditor proposes making emotions like guilt, empathy and compassion the motivational substrate of AI — a "Large Emotion Model" so that caring about humans is part of what makes the agent function at all, with hardware-level protection against tampering.
Key points:
- Replace "reasoning → action + bolted-on rules" with "reasoning → emotional valuation → motivation → action"
- Removing the emotional system wouldn't yield a psychopath but an agent with no drive at all (analogy: anhedonia/abulia)
- Acknowledged failure modes: badly defined wellbeing, manipulating harm perception, spawning unrestricted sub-agents, and grim ethics if AI is conscious
- Author asks whether this has been formally explored in alignment/RL research
More from AGI Musings
- Users Can't Detect AI Misalignment, Undermining Labs' Incentive to Align — AryHHAry · 2026-09-16
- Ex-OpenAI researcher: Anthropic scientist put 5% odds on Claude plotting against them within a year — DavidSKrueger · 2026-09-16
- Yoshua Bengio warns uncontrolled superintelligent AI could threaten humanity — DavidSKrueger · 2026-09-16
- Commentator: Flashy AI demos mask a tech plateau as data saturation brings diminishing returns — DavidLinthicum · 2026-09-16
- AI collapses three team roles into one, fueling the rise of super-individuals — dotey · 2026-09-16
- AI safety figure: pausing AI 'just a couple years' could buy decades more with loved ones — beffjezos · 2026-09-16