Does RL build new representations or just make models flail? LLM researchers debate
doomslide · x · 2026-09-18
voooooogel and doomslide debate what RL actually does to LLMs. voooooogel's intuition: RL both pushes the frontier beyond the model's own understanding (causing flailing and desperation) and builds representations in novel ways, drifting away from human theory-of-mind intuitions — e.g. animacy reversal patterns in 'Claudelish'.
He argues models currently learn search heuristics more than deep understanding. doomslide, working on a persona-selection-style model of LLM behavior, calls the origin of representations a big hole in such frameworks and is cautiously more optimistic about LLMs generating useful new representations.
Related event: Debate: Does RL Create New Representations in LLMs?(2 posts)→
More from Research
- World-SimReady-Home, a multimodal robotics simulation dataset, trends on Hugging Face — Yootta · 2026-09-18
- UHAS demos UHAS visualization: one deformed sphere drives five different dexterous hands — YuXiang_IRVL · 2026-09-18
- Dropping the vector DB from agent tool selection: same recall, cost basically gone — BenefitGrand8752 · 2026-09-18
- Epoch AI launches Benchmark Reviews: only 4 of first 15 benchmarks earn Verified status — xeophon · 2026-09-18
- OpenAI Foundation launches second science program with $125M+ in health data grants — owl_posting · 2026-09-18
- Recursive training collapses LLMs by gen 9; 10% human data halts the damage — alex_verem · 2026-09-18