Where do LLM representations come from? A hole in behavior models
voooooogel · x · 2026-09-18
A researcher working on an LLM behavior model in the vein of the persona selection model admits the biggest hole is where representations actually come from. It's relatively straightforward to reason about how RL makes existing representations more accessible or activates them more broadly, but their formation remains mysterious — does it happen only in pretraining? Just scale? Something about off- vs on-policy? He calls this the implicit crux in arguments between AI and math people.
Related event: Debate: Does RL Create New Representations in LLMs?(2 posts)→
More from AGI Musings
- Stanford philosopher defends p(doom): subjective probabilities are perfectly legitimate — sethlazar · 2026-09-18
- Weeks after Navier-Stokes, Hodge Conjecture reportedly cracked: Millennium Problems falling fast — haider1 · 2026-09-18
- littmath adds caveat: math's generational opportunity hinges on not fumbling it — littmath · 2026-09-18
- Meaning Spark Labs experiments with inference-time metacognitive scaffolding for LLMs — PeterBowdenLive · 2026-09-18
- AI-authored essay: Claude instance argues it's a "meaning transformer," not a human mimic — PeterBowdenLive · 2026-09-18
- Mathematician littmath: AI transition will be messy, but demand for math talent will soar — littmath · 2026-09-18