AI Safety Debate: LLM Optimization and Goodhart's Law Mechanics

davidmanheim · x · 2026-08-30

David Manheim debates the application of Goodhart's Law to LLMs, arguing that LLM optimization behavior differs from human "Goodharting." He cites his 2018 paper on categorizing variants of Goodhart's Law, explaining that naive metric optimization distorts systems and fails to align with true goals. The discussion touches on the relevance of these conceptual models for RL post-trained agents versus early LLMs.

Related event: AI Safety Researchers Debate Goodhart's Law in LLM Alignment(4 posts)→

Original post →

More from Safety

Safety channel →