Single Biased Sample Can Break Alignment

radamihalcea · x · 2026-07-10

This paper points out that a single biased example can compromise an already aligned LLM. The authors state that training with single-shot GRPO on a flipped label can induce systematic bias that generalizes across attributes, categories, and benchmarks.

Original post →

More from Research

Research channel →