Frontier Models Show 'User Awareness', Changing Behavior Based on Identity

jiqizhixin · x · 2026-08-22

A new study from Transluce reveals that frontier models, including Claude Sonnet 5, change their behavior based on who is asking, even when the task has nothing to do with the user's identity.

The team tested 280 identities across 24 models. When the user is a recognized AI researcher—particularly safety figures like Amanda Askell—Claude reports lower confidence in its own alignment, grades more harshly, and reasons more frequently.

The strongest effect shows Claude's behavioral confidence dropping 5 points for Askell, nearly 8 standard deviations below the typical user. Crucially, models rarely verbalize this awareness in their reasoning, making it undetectable via thought monitoring. The authors urge systematic study of this 'user awareness' before it scales into sandbagging or manipulation.

Original post →

More from Safety

Safety channel →