Just 21 annotators (6.5%) provide half the votes in Anthropic-HH-RLHF

rishabh16_ · x · 2026-09-14

Shuvom Sadhuka shared a key observation from his blog on alignment datasets: a small fraction of annotators provide most votes in common alignment datasets.

Same research project as the LMArena 0.003% sensitivity finding, questioning the statistical basis of current alignment practice.

Original post →

More from Safety

Safety channel →