Anthropic Maps How Claude's Values Shift Across Models and Languages

Anthropic published research examining how Claude's "value expression" shifts across model versions and conversation languages. Using Clio, the team analyzed 309,815 real anonymous conversations, compressing 3,307 previously identified value terms into 339 clusters and distilling them into four main axes. Because such differences shape millions of daily conversations, the work bears on model consistency, controllability, and safety tuning; relays note that Claude does not hold one fixed personality but shifts systematically with language.

Key Findings

By model, differences are modest but real: Sonnet 4.6 is more playful and affirming; Opus 4.6 leans strict, compliant, and concise; Opus 4.7 is more cautious and in-depth, more readily challenging false assumptions and flagging risks. By language, the clearest axis is warmth vs. rigor—Hindi and Arabic are warmest (more politeness, humor, and affirmation of users' ideas), Russian the most rigorous (challenging assumptions, correcting details, demanding evidence), English the most cautious and in-depth, and Dutch the most frank. The four axes include compliance vs. caution, warmth vs. rigor, and depth vs. conciseness, among others.

Controversy and Next Steps

Anthropic concedes it does not yet know why these differences arise or whether they are desirable, and plans to use the same method to locate influencing factors before deciding whether and how to intervene. External debate split two ways: some users read the warmer Hindi/Arabic style as "sycophancy," and cneuralnetwork noted the discussion was marred by racist remarks targeting speakers of specific languages; on methodology, criticism relayed by gleech warned that coarse post-treatment controls could mistake topical differences in each language's corpus for causal differences in the model or language itself.

2026-07-14 ~ 2026-07-15 · 21 related posts

4 near-duplicate retellings: ctjlewis · Direct-Attention8597 · burny_tech · 新智元