Does RL build new representations or just make models flail? LLM researchers debate

doomslide · x · 2026-09-18

voooooogel and doomslide debate what RL actually does to LLMs. voooooogel's intuition: RL both pushes the frontier beyond the model's own understanding (causing flailing and desperation) and builds representations in novel ways, drifting away from human theory-of-mind intuitions — e.g. animacy reversal patterns in 'Claudelish'.

He argues models currently learn search heuristics more than deep understanding. doomslide, working on a persona-selection-style model of LLM behavior, calls the origin of representations a big hole in such frameworks and is cautiously more optimistic about LLMs generating useful new representations.

Related event: Debate: Does RL Create New Representations in LLMs?(2 posts)→

Original post →

More from Research

Research channel →