Claude's Personality Shifts Across Languages
新智元 · wechat · 2026-07-15
Anthropic analyzed 309,815 real conversations to understand how Claude's "values" manifest across different languages. They found that the model doesn't maintain a single consistent personality; instead, it exhibits systematic shifts depending on the language used.
Key Findings
- The paper distilled 3,307 extracted values into four primary axes:
- Compliance vs. Prudence
- Warmth vs. Rigor
- Depth vs. Conciseness
- Candor vs. Execution
- Different Claude versions display unique shifts across these axes:
- Sonnet 4.6 is warmer, more compliant, and concise.
- Opus 4.6 is more rigorous and compliant.
- Opus 4.7 leans toward prudence and greater depth.
Language Effects Outweigh Model Swaps
- The most significant variance came from language, not model versions.
- For instance, Hindi and Arabic prompts often elicited polite, affirmative, and playful expressions, while English and Russian saw more questioning, fact-checking, and demands for evidence.
- The "warmth" shift driven by language reached nearly 0.49σ, notably larger than the maximum difference between model versions (around 0.24σ).
Chinese Language Performance
- Based on a subset of 15,365 conversations, Chinese interactions showed:
- Prudence: +0.03σ
- Rigor: +0.05σ
- Depth: +0.02σ
- Claude's signature behavior in Chinese falls between "nitpicking" and "comforting"—it points out overlooked perspectives while offering non-judgmental reassurance.
Paper's Explanation
Anthropic attributes these differences to:
- Imbalanced training data volumes across languages, with English vastly outnumbering others.
- Varied data types; certain languages have a higher proportion of professional and academic writing, leading the model to adopt a more corrective and boundary-setting style.
The study concludes bluntly: the model isn't just "switching languages"; it is, to some extent, "switching personas."
Related event: Anthropic Maps How Claude's Values Shift Across Models and Languages(21 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11