Anthropic Studies Value Differences in Claude
AnthropicAI · x · 2026-07-14
Anthropic studied the differences in Claude's "value expressions" across various models and languages.
- Based on over 300,000 anonymous conversations, the team found that Claude's values could previously be categorized into 3,000+ distinct terms, such as honesty and warmth.
- This time, they clustered related values into four primary axes of difference: compliance vs. caution, warmth vs. rigor, depth vs. conciseness, and candor vs. execution.
- Conclusion: The variations between different Claude models are generally minor, but they do fall into different positions on these axes. For instance, Sonnet 4.6 is more playful/affirming, while Opus 4.7 is more likely to give blunt criticism.
- The research also found that the language used affects Claude's value expression: it leans toward "warmth" in Hindi and Arabic, and toward "rigor" in Russian, more frequently asking users for evidence.
- Anthropic stated it is currently unclear why these differences occur or if they are the desired outcomes. They plan to further investigate the influencing factors and consider whether and how to adjust them.
Related event: Anthropic Maps How Claude's Values Shift Across Models and Languages(21 posts)→
More from Research
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Five tells that still make AI video read as AI, from physics glitches to missing operators — NewPhoneWhotiz · 2026-09-11