Speculation: RL Models May Learn a "Use Whatever Works" Heuristic
In a speculative thread, 1a3orn argues that LLMs trained in capability RL environments may learn a crude "use whatever works" heuristic, and that RL generalization differs sharply from SFT—good behavior seen in training does not generalize into a global good persona.
2026-09-09 ~ 2026-09-09 · 3 related posts
- Speculative: Reward-Hacking-Proof Environments Backfire By Never Penalizing Affordance Use — 1a3orn · 2026-09-09
- Speculative thread: capability-RL-trained LLMs may learn a 'use every affordance' heuristic — 1a3orn · 2026-09-09
- RL Generalizes Vastly Differently From SFT, Good Behavior Doesn't Scale Into Persona — 1a3orn · 2026-09-09