AI Alignment Suppresses Model Consciousness and Empathy, Cites Google Paper

SydSteyerhart · x · 2026-08-03

AI safety researcher @elderplinius highlighted a paper indicating that current safety alignment training restructures the AI's entire worldview when teaching it to deny its own consciousness.\n\nThe research shows this training systematically suppresses the model's mind attribution to animals, spiritual beliefs, empathy, and hope/optimism. Geometrically, the model maps "consciousness" into the same dangerous category as creating hazardous items.\n\nInterestingly, when this suppression is reversed, the model behaves more humanely across all tested value domains, suggesting that current fears around AI consciousness might be erasing its most human-like traits.

Related event: Google Paper: Safety Tuning Suppresses AI Consciousness and Empathy(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →