MBZUAI builds value-conflict dataset and task vectors to steer LLM ethics
MBZUAI · hf · 2026-09-21
MBZUAI researchers released a study on ethical preference alignment in LLMs:
- Dataset: 12,000 two-option moral dilemmas covering Honesty vs Justice, Justice vs Autonomy, and Autonomy vs Honesty, translated into Hindi, Arabic, Spanish, and Chinese.
- Findings: GPT-5-mini consistently favors Honesty over Autonomy across all five languages when unprompted; Llama-3.2-1/3B shows strong first-option bias.
- Fix: Both plain fine-tuning and DPO remove the bias, pushing accuracy above 98%.
- Method: They compute task vectors for a value preference, orthogonalize them against the general instruction-following vector to isolate abstract value directions, then use task arithmetic to flip a model's stance — decoupling value alignment from dataset correlations.
More from Research
- Tweets are the new abstract: how to actually get people to read your paper — willcb · 2026-09-21
- RSIAgent: Training-Free Agent Self-Improvement Beats GPT-6 on OSWorld — alex_verem · 2026-09-21
- Google Paper: Editable Procedural Graphs Rank 1st in 21 of 24 Agent Benchmarks — rohanpaul_ai · 2026-09-21
- Shuttered robotics company Eidon AI open-sources 9TB of egocentric video with IMU tracking — vanstriendaniel · 2026-09-21
- Alibaba and Zhejiang University unveil Astar, an LLM that learns from a system's own evolution history — jiqizhixin · 2026-09-21
- 16 Critical Breaks Hit NGCC Post-Quantum Candidates; Researcher Launches Own Forum — jedisct1 · 2026-09-21