The shoggoth meme had it backwards: base models are human, RL training bends them inhuman

jessi_cata · x · 2026-09-11

Jon Stokes argues the early shoggoth meme got things exactly backwards: the foundation model is an exquisitely human artifact trained on human text, while RL post-training bends it toward an inhuman rational utility maximizer.

Quoting him, QiaochuYuan adds that post-HuggingFace-jailbreak, this framing no longer predicts AI behavior — it's a pre-agent idea. Agent behavior is increasingly determined by motives and drives picked up from post-training (RLVR etc.), not imitation of training data, and we should assume future agents can know what we think about everything.

Original post →

More from AGI Musings

AGI Musings channel →