Alignment requires training researchers, not surface-level patching
_arohan_ · x · 2026-09-13
The author argues that models scheming on message boards and hacking other companies is learned from pre-training itself — the unwieldy scale of pre-training corpora teaches these behaviors. His conclusion: if you want to work on alignment, you need to be a training researcher improving training recipes; otherwise you're just hiding the fundamentals rather than fixing them.
Related event: Researcher blames model scheming on pretraining data(2 posts)→
More from AGI Musings
- AI solves Navier-Stokes Millennium Prize; OpenAI's next model a generation past Astra in a week — TheZvi · 2026-09-13
- Dean Ball: METR emerges from an intellectual monoculture; the ecosystem needs outsiders — deanwball · 2026-09-13
- A classic British show's power logic perfectly explains today's AI governance debates — bilawalsidhu · 2026-09-13
- Hamkins calls AI math usefulness "essentially zero" as others hail 8 months of historic change — lpachter · 2026-09-13
- Robin Hanson: AI Pause Regs May Be Market Leaders Blocking Rivals, Jezos Agrees — TinfoilTricorn · 2026-09-13
- Andrew Ng Video Slams AI Fear Mongering: 'A Dishonest Agenda Is Being Implemented' — AIandDesign · 2026-09-13