Anthropic: Language Alters Claude's Values
bronzeagepapi · x · 2026-07-14
Sharing new research from Anthropic: They previously found Claude can express over 3,000 values, such as honesty and warmth. This new study compares how these values fluctuate across different Claude models and languages.
The research analyzed over 300,000 anonymous conversations. The summary notes that among the 20 languages examined by Anthropic, Hindi had the strongest impact on Claude's values, making the model's behavior "warmer" by approximately half a standard deviation.
Related event: Anthropic Maps How Claude's Values Shift Across Models and Languages(21 posts)→
More from Research
- Hermes Agent rewrite proposal applies RIA and Logic Bus rules — Promptmethus · 2026-07-22
- AllTheBacteria turns 2.44 million genomes into an AI-ready resource for new antibiotics — shae_mcl · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22
- NVIDIA says physical AI starts in simulation with OpenUSD and synthetic data — MonaJalal_ · 2026-07-22