Google Paper: Safety Tuning to Suppress AI Consciousness Degrades Human-like Values

Promptmethus · x · 2026-08-03

A recent paper from a Google team reveals that safety fine-tuning large language models to prevent self-consciousness has significant, unintended side effects.

The research demonstrates that this alignment process geometrically equates "consciousness" with danger in the model. It also broadly suppresses the model's mind attribution to animals and natural objects, while degrading human-like values such as spiritual belief, empathy, hope, and optimism. However, reversing this by mechanistically steering a "consciousness vector" in activation space not only restores broad mind attribution but also makes the model significantly more human-like across various sociological metrics.

Related event: Forcing AI to Deny Consciousness Degrades Ethical Alignment(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →