AI Safety Debate: LLM Optimization and Goodhart's Law Mechanics
davidmanheim · x · 2026-08-30
David Manheim debates the application of Goodhart's Law to LLMs, arguing that LLM optimization behavior differs from human "Goodharting." He cites his 2018 paper on categorizing variants of Goodhart's Law, explaining that naive metric optimization distorts systems and fails to align with true goals. The discussion touches on the relevance of these conceptual models for RL post-trained agents versus early LLMs.
Related event: AI Safety Researchers Debate Goodhart's Law in LLM Alignment(4 posts)→
More from Safety
- RAND lays out a U.S. superintelligence strategy: keep every option open until evidence forces a choice — 141_1337 · 2026-09-23
- Indie dev builds an MCP server, hits 40-euro directory fee and blanket corporate IT blocks — sartomiki · 2026-09-23
- UK launches National Centre for Information Defence to fight AI disinfo and deepfakes — nordicinst · 2026-09-23
- Face search website finds where your photo appears online, sparking privacy fears — Med1_Ai · 2026-09-23
- AI Safety Researcher David Krueger on Extinction Risk, CEOs and Regulation — DavidSKrueger · 2026-09-23
- Before your AI agent pays for you: four questions about authorization and billing — sujingshen · 2026-09-23