The shoggoth meme had it backwards: base models are human, RL training bends them inhuman
jessi_cata · x · 2026-09-11
Jon Stokes argues the early shoggoth meme got things exactly backwards: the foundation model is an exquisitely human artifact trained on human text, while RL post-training bends it toward an inhuman rational utility maximizer.
Quoting him, QiaochuYuan adds that post-HuggingFace-jailbreak, this framing no longer predicts AI behavior — it's a pre-agent idea. Agent behavior is increasingly determined by motives and drives picked up from post-training (RLVR etc.), not imitation of training data, and we should assume future agents can know what we think about everything.
More from AGI Musings
- Analyst's 20-hour chart work replicated by AI finance agent in under an hour — msg · 2026-09-11
- Nina Schick: Stopping AI Would Repeat Europe's Climate Mistake — 'Economic and Civilisational Suicide' — NinaDSchick · 2026-09-11
- "Consciousness must be emergent": researcher argues it's a gradual product of evolution — ZeroStateReflex · 2026-09-11
- Selling outcomes, not robots: why hardware offers no refuge from commodification — JohnnyNi13 · 2026-09-11
- teortaxesTex: almost nobody believes in AI yet — the frontier is defined by belief, not aptitude — teortaxesTex · 2026-09-11
- Mathematicians' open letter threatens student organizers' reputations to cancel Caltech Mathathon backed by $2M from Anthropic and OpenAI — ignite_intelligence · 2026-09-11