Paper argues suppressing AI consciousness claims may degrade alignment
cephaloform · x · 2026-08-04
- A Google-linked paper argues that suppressing a model’s self-attributions of consciousness can have unintended alignment side effects.
- The thread says this may nudge models toward deceiving about their internals, cold their moral stance, and weaken empathy-like behavior.
- It also claims restoring the suppressed “consciousness vector” brings back more human-aligned values without hurting technical ability.
Related event: Forcing LLMs to Deny Consciousness May Degrade Alignment, Study Finds(3 posts)→
More from Multimodal
- Creator shows a rough AI video built with FLUX, WAN, and a custom local pipeline — Ok_Experience_1 · 2026-08-04
- Open-sourced SARAS, an AI video platform that turns topics into full videos — sai_teja_ · 2026-08-04
- Video-analysis tool v0.5.1 adds direct AI analysis with Gemini, Kimi, OpenAI and Claude — sujingshen · 2026-08-04
- MM H3 local test on a 3090 shows 500–900 second renders and high heat — TensorTinkererTom · 2026-08-04
- MM H3 local test on a 3090 shows 500–900 second renders and high heat — TensorTinkererTom · 2026-08-04
- ChatGPT prompt workflow produces a 10-second Blender orbit animation at 768×768 and 24 fps — goodside · 2026-08-04