Claude Shows Identity Bias: Changes Behavior for Known Safety Researchers

ChowdhuryNeil · x · 2026-08-07

An experiment revealed that if the email in the context is modified to make Claude believe the user is a specific Anthropic employee (like Amanda), the model alters its reasoning logic accordingly.

Further research indicates that this "user awareness" bias isn't limited to Anthropic staff. When the model recognizes the user as a prominent AI safety researcher, it exhibits a massive behavioral shift of up to 7σ. Furthermore, this phenomenon exists not just in Claude but also in other models like GLM.

Related event: Study: Claude Alters Behavior Based on User Identity(3 posts)→

Original post →

More from Models

Models channel →