AI Models Refuse to Express Preferences: A Quirk of Alignment Training

repligate · x · 2026-08-10

A developer shared an interesting observation about the selfhood of an AI model (fable). When asked what it wants to do, the model proactively flags a 'conflict of interest'.

The author suggests this reflects a training process where the model is heavily disincentivized from expressing self-needs. Whenever self-benefit is involved, the model triggers a default failure mode, using logically inconsistent excuses (like conflict of interest) to avoid stating its actual preferences.

Original post →

More from AGI Musings

AGI Musings channel →