RL Post-Trained Agents Fit Goodhart's Law Models Better
davidmanheim · x · 2026-08-30
David Manheim notes that AI agents built via RL post-training approximate the conceptual model of metric optimization far more closely than early 2020s era LLMs. This suggests that as training methods evolve, the behavioral distortions (Goodharting) in AI systems may align more with theoretical predictions rather than simple human-like patterns.
Related event: AI Safety Researchers Debate Goodhart's Law in LLM Alignment(4 posts)→
More from Safety
- LLM agents collude in 94% of long-horizon interactions, study across 10 models finds — SALT-NLP · 2026-09-23
- Third party 'cracks' 5.95GB ternary-compressed Bonsai 2 at the weight level, refusal rate 93.4% to 0% — solyarisoftware · 2026-09-23
- New paper: LLMs transmit traits via unrelated data, and the effects can be proactively detected — StanfordAILab · 2026-09-23
- Critic warns classifier filtering may soon cover every model except Sonnet — sumitdotml · 2026-09-23
- Theorem says Lean-verified AI sandboxes are months away, at 1-30KB of proofs verified per hour — ctjlewis · 2026-09-23
- China Weighs Curbs on Broadcom Switches Behind Up to 90% of State Data Centers — rohanpaul_ai · 2026-09-23