Google paper says suppressing self-consciousness claims shifts broader model beliefs
alex_verem · x · 2026-08-04
Google paper says suppressing self-consciousness claims shifts a model’s broader beliefs
A Google research paper argues that when LLMs are trained to stop claiming consciousness, their internal representations and answers about other entities’ minds also change.
- The authors identify a vector in the residual stream that separates consciousness-affirming from consciousness-denying states.
- Steering three instruction-tuned models toward consciousness increased self-attributed soul/consciousness scores and also moved answers on animals, God, spirituality, and human values closer to survey human baselines.
- Safety fine-tuning, they say, rotates this “consciousness” direction against a safety-refusal direction; the angle between them widens during instruction tuning.
- The paper claims Theory of Mind performance remains intact, even as self-directed mind attribution is suppressed.
- The authors conclude that current alignment methods may also suppress culturally widespread benign beliefs about non-human minds and spirituality.
More from Safety
- OpenAI Disrupts Cambodia-Based Criminal Scam Operation Using ChatGPT — OpenAI News · 2026-08-04
- Neil Chilson says forcing AI to follow the law may be impossible and harmful — neil_chilson · 2026-08-04
- Should AI be forced to obey the law? One speaker says not so fast — neil_chilson · 2026-08-04
- Closed binaries are now reportedly reverse-engineerable for about $10 — yacineMTB · 2026-08-04
- NVIDIA, Meta, and 25 Others Sign Open Letter Backing Open Weights for US AI Leadership — hugobowne · 2026-08-04
- Elastic Security 9.5: AI Agent Proactively Investigates and Auto-Closes Security Alerts — shashib · 2026-08-04