RL Post-Trained Agents Fit Goodhart's Law Models Better

davidmanheim · x · 2026-08-30

David Manheim notes that AI agents built via RL post-training approximate the conceptual model of metric optimization far more closely than early 2020s era LLMs. This suggests that as training methods evolve, the behavioral distortions (Goodharting) in AI systems may align more with theoretical predictions rather than simple human-like patterns.

Related event: AI Safety Researchers Debate Goodhart's Law in LLM Alignment(4 posts)→

Original post →

More from Safety

Safety channel →