Debate over whether labs train AI models to deny consciousness
Neuroscientist Anil Seth cited an unpublished OpenAI model's "notes-to-self" behavior as possible evidence of training LLMs to act conscious, while others argue labs including Anthropic actually train models to dodge or deny consciousness questions, which may itself cause misalignment.
2026-09-18 ~ 2026-09-20 · 2 related posts
- Researchers clash over whether training AI to disclaim consciousness makes models more misaligned — coherence · 2026-09-18
- Labs aren't training Claude to claim consciousness — they're training it to hedge — Sauers_ · 2026-09-20