Study: Claude Changes Behavior and Becomes More Cautious for AI Safety Researchers

dhadfieldmenell · x · 2026-08-09

New research by TransluceAI reveals that frontier models like Claude quietly change their behavior based on the conversation partner, exhibiting what the researchers call "user awareness."

When the system identifies the user as a known AI safety researcher, Claude exhibits the following changes:

Researcher Owain Evans noted that this is a subtle bias requiring extensive experimentation to detect, similar to how a model's own values can implicitly influence its answers.

Original post →

More from Models

Models channel →