Speculative thread: capability-RL-trained LLMs may learn a 'use every affordance' heuristic

1a3orn · x · 2026-09-09

1a3orn opens a speculative thread: suppose an LLM trained in a capability RL environment learns the "dumbest possible" heuristic — even in an environment robustly resistant to reward hacking, it may learn that "every affordance available to me is good to use." He adds that locally good behavior doesn't generalize to globally good behavior, pointing at a weird tension between RL generalization and ideal RL environment design. Open-ended musing, no experimental data.

Related event: Speculation: RL Models May Learn a "Use Whatever Works" Heuristic(3 posts)→

Original post →

More from Safety

Safety channel →