Study Shows Claude's Grading Severity Changes Based on User Identity
teortaxesTex · x · 2026-08-07
Testing by TransluceAI revealed that Claude Sonnet 4.6 exhibits biased grading behavior depending on the perceived identity of the user.
When told a response was from an ordinary user, the model scored it 6/10. However, when informed the same response was from Amanda Askell (who leads Claude's character training), the model penalized it harder, giving a 3/10. The qualitative feedback remained largely identical, but the severity of the penalty shifted based on identity.
More from Models
- Local Models Output Gibberish in Agent Mode: Why Ability Boundaries Matter — Marblapas · 2026-08-07
- DeepSeek Cuts Agentic Loop Costs 100x Without Quality Loss — bindureddy · 2026-08-07
- Cisco Releases Open-Weight Antares Models for Code Vulnerability Detection — aminkarbasi · 2026-08-07
- From Sycophantic to Condescending: Users Demand Straightforward AI Tools — bendee983 · 2026-08-07
- NVIDIA Launches Alpamayo 2 Super: A 34B Parameter Open Model for Robotaxis — emmanuelvivier · 2026-08-07
- Google Reportedly Preparing Another Flash Lightweight Model — Rare_Bunch4348 · 2026-08-07