Where do LLM representations come from? A hole in behavior models

voooooogel · x · 2026-09-18

A researcher working on an LLM behavior model in the vein of the persona selection model admits the biggest hole is where representations actually come from. It's relatively straightforward to reason about how RL makes existing representations more accessible or activates them more broadly, but their formation remains mysterious — does it happen only in pretraining? Just scale? Something about off- vs on-policy? He calls this the implicit crux in arguments between AI and math people.

Related event: Debate: Does RL Create New Representations in LLMs?(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →