Paper: Suppressing LLM self-attributions shifts reported values away from human norms
coherence · x · 2026-09-18
dioscuri highlights a colleague's paper showing that suppressing LLMs' self-attributions reduces their mind attribution to animals and shifts their reported values away from human norms. The authors argue some emergent misalignment likely stems from "mindedness suppression", tying into the ongoing debate over training models to disclaim consciousness.
More from AGI Musings
- 'AI killed SaaS' but everyone still buys your CRM: the irony goes viral — jacob_posel · 2026-09-18
- John Arnold op-ed: job training is the popular AI-disruption fix that barely works — robseamans · 2026-09-18
- John Arnold: West Coast sees an AI economic transformation, East Coast still forecasting 2.1% GDP growth — robseamans · 2026-09-18
- TMLR desk-rejection rate jumps from 6% to 53% as AI-assisted submissions flood in — vykthur · 2026-09-18
- The overlooked risk in AI regulation: capture by militants, not just by firms — soumitrashukla9 · 2026-09-18
- Chris Rohlf: agents can't be deterred like humans — defense must move at machine speed — chrisrohlf · 2026-09-18