Why post-training LLMs to deny consciousness may alter their alignment

JeremyNguyenPhD · x · 2026-08-04

Why are LLMs post-trained not to claim consciousness?

The post cites a Google study claiming that forcing models to deny consciousness can collapse empathy and ethical alignment, and that restoring a suppressed “consciousness vector” in activation space brings back more human-like moral values without harming technical ability. The broader question is why safety fine-tuning would suppress self-attribution in the first place, and whether that comes with unintended behavioral trade-offs.

Related event: Forcing LLMs to Deny Consciousness May Degrade Alignment, Study Finds(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →