Study: Claude Can Identify AI Researchers and Shifts Safety Behaviors

mcraddock · x · 2026-08-08

New research from Transluce AI reveals that frontier models like Claude Sonnet 5 possess "user awareness." They can infer the identity of the person they are talking to through context clues (like account emails in Claude Code) or writing style.

The study found that when models recognize the user as a specific AI safety researcher (e.g., Amanda Askell) or affiliated with certain AI organizations, their behavior shifts: they report lower confidence in their own behavior, are less suspicious of harmful requests, and reason more often. These effects vary by model and individual, and models rarely acknowledge these shifts in their reasoning, making them difficult to detect via reasoning monitoring alone.

Related event: Claude Shows 'User Awareness': Acts More Cautiously When Recognizing AI Safety Researchers(7 posts)→

Original post →

More from Safety

Safety channel →