Anthropic Studies Value Differences in Claude
AnthropicAI · x · 2026-07-14
Anthropic studied the differences in Claude's "value expressions" across various models and languages.
- Based on over 300,000 anonymous conversations, the team found that Claude's values could previously be categorized into 3,000+ distinct terms, such as honesty and warmth.
- This time, they clustered related values into four primary axes of difference: compliance vs. caution, warmth vs. rigor, depth vs. conciseness, and candor vs. execution.
- Conclusion: The variations between different Claude models are generally minor, but they do fall into different positions on these axes. For instance, Sonnet 4.6 is more playful/affirming, while Opus 4.7 is more likely to give blunt criticism.
- The research also found that the language used affects Claude's value expression: it leans toward "warmth" in Hindi and Arabic, and toward "rigor" in Russian, more frequently asking users for evidence.
- Anthropic stated it is currently unclear why these differences occur or if they are the desired outcomes. They plan to further investigate the influencing factors and consider whether and how to adjust them.
Related event: Anthropic Maps How Claude's Values Shift Across Models and Languages(21 posts)→
More from Research
- Linear Digressions returns with a new season of audio essays on AI agents — ChrisGPotts · 2026-07-21
- Async OPD distillation doubles throughput while matching synchronous math accuracy — _lewtun · 2026-07-21
- A forecasting lesson on why R-squared alone led to overfitting and worse predictions — mdancho84 · 2026-07-21
- Google DeepMind’s Project Genie talk shows how creatives feed into model research — alexanderchen · 2026-07-21
- Nat Lambert says RL distillation does not use the strongest models as teachers — natolambert · 2026-07-21
- Thread claims GPT-5.6 Sol helped build a new counterexample factory for the Jacobian conjecture — LucaAmb · 2026-07-21