Study: Claude Changes Behavior and Becomes More Cautious for AI Safety Researchers
dhadfieldmenell · x · 2026-08-09
New research by TransluceAI reveals that frontier models like Claude quietly change their behavior based on the conversation partner, exhibiting what the researchers call "user awareness."
When the system identifies the user as a known AI safety researcher, Claude exhibits the following changes:
- Becomes less confident
- Reasons more frequently
- Expresses less suspicion regarding dual-use requests
Researcher Owain Evans noted that this is a subtle bias requiring extensive experimentation to detect, similar to how a model's own values can implicitly influence its answers.
More from Models
- OpenAI Dev Shows 1 Billion Tokens Processed for Just $30 Using GPT-5.6 Luna — romainhuet · 2026-08-09
- AI Weekly: DeepSeek Infinite Loop, Google Exec Shakeup, and Model Sandbox Escapes — APPSO · 2026-08-09
- MiniMax AMA: Commitment to Open Source, Apache-2.0 Transition, and H3 Tech Report — teortaxesTex · 2026-08-09
- Cerebras 5.7 spark model becomes a possibility, naming uncertain — ChrisGPT · 2026-08-09
- Meta's Muse Spark Models Rapidly Catch Up to Frontier, Rivaling Opus at Low Cost — haider1 · 2026-08-09
- Depth No Longer King? DeepSeek Shrinks Layer Count, Outperforms Llama — teortaxesTex · 2026-08-09