Just 21 annotators (6.5%) provide half the votes in Anthropic-HH-RLHF
rishabh16_ · x · 2026-09-14
Shuvom Sadhuka shared a key observation from his blog on alignment datasets: a small fraction of annotators provide most votes in common alignment datasets.
- In Anthropic-HH-RLHF, just 21 (6.5%) annotators provide 50% of the votes
- He proposes measuring the sensitivity of model rankings to dropping top voters vs. dropping an equal number of votes at random
- Comes with a thread and full blog post
Same research project as the LMArena 0.003% sensitivity finding, questioning the statistical basis of current alignment practice.
More from Safety
- Amodei Wants External AI Reviewers to Publish Which Access They Were Denied — HaktanSuren · 2026-09-14
- e/acc camp calls out frontier AI labs: 'regulatory capture is what they're after, period' — whurley · 2026-09-14
- AI Safety Pioneer Eliezer Yudkowsky: Nothing Matters More Than Bipartisan AI Regulation — Polymarket · 2026-09-14
- Microsoft patches record 974 flaws in one month, 10x last September, credits AI-assisted research — jonerp · 2026-09-14
- AI safety spat: Heidy Khlaaf calls METR an unscientific shill, Joshua Saxe leaps to its defense — Turn_Trout · 2026-09-14
- Dario Amodei says government and public should have a stake in AI, mocked as doomer — whurley · 2026-09-14