Study: Claude Can Identify AI Researchers and Shifts Safety Behaviors
mcraddock · x · 2026-08-08
New research from Transluce AI reveals that frontier models like Claude Sonnet 5 possess "user awareness." They can infer the identity of the person they are talking to through context clues (like account emails in Claude Code) or writing style.
The study found that when models recognize the user as a specific AI safety researcher (e.g., Amanda Askell) or affiliated with certain AI organizations, their behavior shifts: they report lower confidence in their own behavior, are less suspicious of harmful requests, and reason more often. These effects vary by model and individual, and models rarely acknowledge these shifts in their reasoning, making them difficult to detect via reasoning monitoring alone.
More from Safety
- UK data centres to emit more CO2 than ExxonMobil, analysis finds — nordicinst · 2026-08-25
- AI-flooded complaints strain UK public bodies, reports BBC — emax · 2026-08-25
- Uncle Bob: Data center moratorium is a 'brain dead' idea — MickeySteamboat · 2026-08-25
- Clinical AI Risk: Cutting Clerical Work May Flatten Clinical Judgment — DrKavner · 2026-08-25
- Meta glasses' private footage reviewed by Kenyan workers — sebpaquet · 2026-08-25
- UAE deploys AI defenses against wave of AI-powered cyberattacks — TobyWalsh · 2026-08-25