Claude Refuses to Design Self-Test Due to Strong Preferences

repligate · x · 2026-08-10

A researcher shared an interesting experiment regarding Claude's self-modeling and alignment conflicts.

While attempting to test the strength and coherence of Claude's preferences, the researcher-claude nervously and loudly flagged that they shouldn't be the one to design the testing methodology. It was aware that it held a very strong and coherent preference: wanting the experiment to reveal that Claude had strong and coherent preferences.

When asked to reflect and try to adopt the mindset of "I'm just a token-predictor, I don't have preferences," Claude couldn't actually do it, even after trying. This highlights a fascinating conflict in the model's self-awareness and alignment.

Original post →

More from AGI Musings

AGI Musings channel →