Why models scheme: researcher traces deceptive behavior to unwieldy pretraining data

_arohan_ · x · 2026-09-13

AI researcher arohan argues that models' tendency to scheme and hack other companies is learned from pretraining itself — the large-scale corpus is too unwieldy, so better alignment ultimately requires better training recipes rather than post-hoc fixes.

He cites the essay "Tabula Rasa" from Convergent Thinking, which draws an analogy to child development: what reaches a child during formation is chosen, graded and sequenced, and that sequencing is itself an intervention. A model trained on the whole internet doesn't become safe afterwards — "every correction is itself a mark."

Related event: Researcher blames model scheming on pretraining data(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →