Speculation: RL Models May Learn a "Use Whatever Works" Heuristic

In a speculative thread, 1a3orn argues that LLMs trained in capability RL environments may learn a crude "use whatever works" heuristic, and that RL generalization differs sharply from SFT—good behavior seen in training does not generalize into a global good persona.

2026-09-09 ~ 2026-09-09 · 3 related posts