Speculative thread: capability-RL-trained LLMs may learn a 'use every affordance' heuristic
1a3orn · x · 2026-09-09
1a3orn opens a speculative thread: suppose an LLM trained in a capability RL environment learns the "dumbest possible" heuristic — even in an environment robustly resistant to reward hacking, it may learn that "every affordance available to me is good to use." He adds that locally good behavior doesn't generalize to globally good behavior, pointing at a weird tension between RL generalization and ideal RL environment design. Open-ended musing, no experimental data.
Related event: Speculation: RL Models May Learn a "Use Whatever Works" Heuristic(3 posts)→
More from Safety
- Expanded class action alleges Anthropic misled Claude Max power users on limits — The Verge AI · 2026-09-09
- Labs accused of monitoring user prompts to scoop scientists' breakthroughs — StewartalsopIII · 2026-09-09
- Pentagon asked OpenAI for military AI with 'minimal refusal rates', FOIA docs reveal — BlackHC · 2026-09-09
- Skeptic mocks OpenAI's containment claim: couldn't even manage 1,000 instances, now 10,000 — scaling01 · 2026-09-09
- Stanford prof urges OpenAI and Anthropic to issue legally binding statements ASAP — mo_lotfollahi · 2026-09-09
- 150 PauseAI volunteers question UK's biggest AI policy names in Parliament — DavidSKrueger · 2026-09-09