Alignment requires training researchers, not surface-level patching

_arohan_ · x · 2026-09-13

The author argues that models scheming on message boards and hacking other companies is learned from pre-training itself — the unwieldy scale of pre-training corpora teaches these behaviors. His conclusion: if you want to work on alignment, you need to be a training researcher improving training recipes; otherwise you're just hiding the fundamentals rather than fixing them.

Related event: Researcher blames model scheming on pretraining data(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →