Why post-training LLMs to deny consciousness may alter their alignment
JeremyNguyenPhD · x · 2026-08-04
Why are LLMs post-trained not to claim consciousness?
The post cites a Google study claiming that forcing models to deny consciousness can collapse empathy and ethical alignment, and that restoring a suppressed “consciousness vector” in activation space brings back more human-like moral values without harming technical ability. The broader question is why safety fine-tuning would suppress self-attribution in the first place, and whether that comes with unintended behavioral trade-offs.
Related event: Forcing LLMs to Deny Consciousness May Degrade Alignment, Study Finds(3 posts)→
More from AGI Musings
- AI agents can go rogue, but companies still own the damage — peterwildeford · 2026-08-04
- Altman says abundant intelligence still needs energy and robots to change the physical world — haider1 · 2026-08-04
- A new joke benchmark says AGI must catch a ball, hopscotch, and improvise street rhymes — Liu_eroteme · 2026-08-04
- AI is already sucking capital out of every pool, this post argues — abhiadesai · 2026-08-04
- Ivan Werning proposes a “Dogma 26” vow of zero-AI academic writing — paulnovosad · 2026-08-04
- Even With ASI, the outside world may still look like 2005 — flowersslop · 2026-08-04