Study: Achieving LLM Cultural Alignment via Feature Steering

simi_97k · x · 2026-07-12

This study (COLM 2026) investigated LLM cultural biases in open-ended questions, finding that default outputs heavily favor US/UK cultures (around 60%).

The research team used Sparse Autoencoders (SAE) to extract features and discovered that model parameters actually contain abundant multicultural knowledge, which is suppressed by standard training. By applying the CuE method for feature steering during inference, the model's cultural bias can be altered transparently and with low data costs to generate culture-specific content. This method outperforms explicit prompting in cultural fidelity and can uncover long-tail cultural concepts. However, it currently relies on SAEs trained for specific models and struggles to dynamically find the optimal steering strength during inference.

Related event: Steering LLMs for Culturally Localized Generation(5 posts)→

Original post →

More from Research

Research channel →