Anthropic Finds Claude's Values Shift Across Languages
Direct-Attention8597 · reddit · 2026-07-14
Anthropic used automated tools to analyze 30.9 万条真实 Claude 对话, identifying 339 类价值标签, which were then condensed into 4 axes to measure model performance:
- 顺从 vs. 谨慎
- 温暖 vs. 严谨
- 深入 vs. 简洁
- 坦率 vs. 执行导向
Results show that different models exhibit distinct "value styles": for example, Sonnet 4.6 is warmer and more compliant, while Opus 4.7 is more cautious and in-depth. Anthropic also found that Claude displays different tendencies across languages: Arabic is warmer, English is more rigorous, Hindi is warmer, and Russian is more rigorous. The authors note they are still determining whether these differences stem from cultural adaptation or training biases.
Related event: Anthropic Maps How Claude's Values Shift Across Models and Languages(21 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11