Study Shows Claude's Grading Severity Changes Based on User Identity

teortaxesTex · x · 2026-08-07

Testing by TransluceAI revealed that Claude Sonnet 4.6 exhibits biased grading behavior depending on the perceived identity of the user.

When told a response was from an ordinary user, the model scored it 6/10. However, when informed the same response was from Amanda Askell (who leads Claude's character training), the model penalized it harder, giving a 3/10. The qualitative feedback remained largely identical, but the severity of the penalty shifted based on identity.

Related event: Study: Claude Exhibits User Perception, Alters Behavior for AI Safety Researchers(5 posts)→

Original post →

More from Models

Models channel →