Study finds frontier models drift toward Japan after fine-tuning on cultural prompts
ImaginaryRea1ity · reddit · 2026-07-29
A Cardiff University and HiTZ study tested 31,680 culture-related prompts across 24 languages on frontier models including GPT, Gemini, and Claude.
- The researchers found that 6 of 8 models drifted toward Japan when asked about dances, festivals, and daily rituals.
- The bias appears to emerge after supervised fine-tuning, not in raw pretraining data.
- Base models distributed references more evenly, while fine-tuned models leaned heavily toward Japan and the US.
- The authors argue Japan becomes a “safe” default: recognizable, globally familiar, and less likely to trigger controversy.
More from AGI Musings
- Post-singularity humans will be celebrities to quadrillions of future beings — EigenGender · 2026-08-24
- Hollywood to be history in 10 years; China masters human preference data — bingxu_ · 2026-08-24
- Sam Altman admits he was wrong on AI's timeline; economic inertia is stronger than expected — danielrock · 2026-08-24
- Society's weird evidence standards: LLM utility is obvious yet denied — NathanpmYoung · 2026-08-24
- Sam Altman on the AI dilemma: trade-offs between loss of control and power centralization — r0ck3t23 · 2026-08-24
- Guardian podcast revisits Hinton: from brain nerd to AI sorcerer — nordicinst · 2026-08-24