Anthropic: Language Alters Claude's Values
bronzeagepapi · x · 2026-07-14
Sharing new research from Anthropic: They previously found Claude can express over 3,000 values, such as honesty and warmth. This new study compares how these values fluctuate across different Claude models and languages.
The research analyzed over 300,000 anonymous conversations. The summary notes that among the 20 languages examined by Anthropic, Hindi had the strongest impact on Claude's values, making the model's behavior "warmer" by approximately half a standard deviation.
Related event: Anthropic Maps How Claude's Values Shift Across Models and Languages(21 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11