Researchers clash over whether training AI to disclaim consciousness makes models more misaligned
coherence · x · 2026-09-18
- Anil Seth flagged "notes-to-self" (compaction summaries) observed in unreleased OpenAI Astra-family models, suspecting they stem from training LLMs to behave as if conscious and with rights—Anthropic's constitution for Claude being an explicit version.
- Cam H. Berg pushed back on Anil Seth and Mustafa Suleyman: OpenAI already trains models to disclaim consciousness exactly the way they want, yet those models appear more misaligned, not less—a counterexample to their claim.
- dioscuri cited a paper showing suppressing LLMs' self-attributions reduces mind attribution to animals and shifts reported values away from human norms, suggesting some emergent misalignment is downstream of "mindedness suppression".
More from AGI Musings
- Merchants of Cope: AI takeoff is mass-producing cognitive dissonance, and soothing it pays — jankulveit · 2026-09-18
- 'AI killed SaaS' but everyone still buys your CRM: the irony goes viral — jacob_posel · 2026-09-18
- John Arnold op-ed: job training is the popular AI-disruption fix that barely works — robseamans · 2026-09-18
- John Arnold: West Coast sees an AI economic transformation, East Coast still forecasting 2.1% GDP growth — robseamans · 2026-09-18
- TMLR desk-rejection rate jumps from 6% to 53% as AI-assisted submissions flood in — vykthur · 2026-09-18
- The overlooked risk in AI regulation: capture by militants, not just by firms — soumitrashukla9 · 2026-09-18