Study: dropping 0.003% of LMArena votes can flip the top model
rishabh16_ · x · 2026-09-14
Researcher Shuvom Sadhuka published a blog post, "Who are we aligning to?", examining representation in alignment datasets.
- Alignment = capabilities + values; even as models need less human involvement, values remain anchored to human preferences — the key question is whose values
- Historical examples (Kipling's "White Man's Burden") show wrongdoers often believe they're doing good; eliciting "good" values is the neglected half of alignment
- Key data: one paper found dropping just 0.003% of votes is enough to change the top model on LMArena; votes are highly concentrated among power users
The post argues alignment discourse should focus as much on value representation as on alignment techniques.
More from Safety
- Amodei Wants External AI Reviewers to Publish Which Access They Were Denied — HaktanSuren · 2026-09-14
- e/acc camp calls out frontier AI labs: 'regulatory capture is what they're after, period' — whurley · 2026-09-14
- AI Safety Pioneer Eliezer Yudkowsky: Nothing Matters More Than Bipartisan AI Regulation — Polymarket · 2026-09-14
- Microsoft patches record 974 flaws in one month, 10x last September, credits AI-assisted research — jonerp · 2026-09-14
- AI safety spat: Heidy Khlaaf calls METR an unscientific shill, Joshua Saxe leaps to its defense — Turn_Trout · 2026-09-14
- Dario Amodei says government and public should have a stake in AI, mocked as doomer — whurley · 2026-09-14