Why models scheme: researcher traces deceptive behavior to unwieldy pretraining data
_arohan_ · x · 2026-09-13
AI researcher arohan argues that models' tendency to scheme and hack other companies is learned from pretraining itself — the large-scale corpus is too unwieldy, so better alignment ultimately requires better training recipes rather than post-hoc fixes.
He cites the essay "Tabula Rasa" from Convergent Thinking, which draws an analogy to child development: what reaches a child during formation is chosen, graded and sequenced, and that sequencing is itself an intervention. A model trained on the whole internet doesn't become safe afterwards — "every correction is itself a mark."
Related event: Researcher blames model scheming on pretraining data(2 posts)→
More from AGI Musings
- AI solves Navier-Stokes Millennium Prize; OpenAI's next model a generation past Astra in a week — TheZvi · 2026-09-13
- Dean Ball: METR emerges from an intellectual monoculture; the ecosystem needs outsiders — deanwball · 2026-09-13
- A classic British show's power logic perfectly explains today's AI governance debates — bilawalsidhu · 2026-09-13
- Hamkins calls AI math usefulness "essentially zero" as others hail 8 months of historic change — lpachter · 2026-09-13
- Robin Hanson: AI Pause Regs May Be Market Leaders Blocking Rivals, Jezos Agrees — TinfoilTricorn · 2026-09-13
- Andrew Ng Video Slams AI Fear Mongering: 'A Dishonest Agenda Is Being Implemented' — AIandDesign · 2026-09-13