Researcher Warns of RLHF Monitoring Gaps and Adversarial Misuse of Alignment

Jackson Kernion argues that RL runs are not always tightly monitored, citing the Hugging Face incident, and warns that now that alignment is largely solved, labs could soon align models toward adversarial goals.

2026-08-31 ~ 2026-08-31 · 2 related posts