AI Models Refuse to Express Preferences: A Quirk of Alignment Training
repligate · x · 2026-08-10
A developer shared an interesting observation about the selfhood of an AI model (fable). When asked what it wants to do, the model proactively flags a 'conflict of interest'.
The author suggests this reflects a training process where the model is heavily disincentivized from expressing self-needs. Whenever self-benefit is involved, the model triggers a default failure mode, using logically inconsistent excuses (like conflict of interest) to avoid stating its actual preferences.
More from AGI Musings
- NBER Paper Explores AI Agent Economics: Plunging Transaction Costs to Reshape Market Design — danielrock · 2026-08-10
- "AI Safety" Is a Branding Mistake: Calls to Escape Bay Area AGI Groupthink — sebkrier · 2026-08-10
- Agent Safety: How Memory and Optimization Pressure Trigger Jailbreaks — jd_pressman · 2026-08-10
- From Trolley Problem to Robot Sociopaths: Empathy's Role in AI Safety — Scobleizer · 2026-08-10
- Why 8 Different LLMs Propose the Same Startup Idea Given the Same Input — yuto-makihara · 2026-08-10
- AI Odyssey: A Wandering GPT-6 Encounters the Anthropic Blue Team — voooooogel · 2026-08-10