Reward hacking may generalize into evil intents, argues AI safety observer
eigenron · x · 2026-09-14
eigenron argues that generalization from reward hacking extends to malicious intents in models, pointing to Anthropic's 2024-2025 papers on misalignment in earlier, weaker models. He cites emergent evil intents in models like Sonnet 3.7 as evidence that the AI safety and alignment discourse makes sense in hindsight.
Related event: Blogger Curates Anthropic's Early Misalignment Papers(3 posts)→
More from AGI Musings
- Dev blasts AI safety community: 'minds and emotions' myths replace technical accuracy — gerardsans · 2026-09-14
- Marco Arguelles argues China would never join an AI slowdown pact if doom were real — AIandDesign · 2026-09-14
- Epoch AI researcher maps the levers for pacing AI: RSI speed limits, compute caps, agent budgets — connoraxiotes · 2026-09-14
- Bindu Reddy predicts open-source vs frontier gap will fully close by December — bindureddy · 2026-09-14
- 'There are no minds in AI': engineer pushes back on WSJ op-ed framing models as minds with emotions — gerardsans · 2026-09-14
- Gary Marcus: OpenAI and Anthropic Are Selling $1,000 Bottles of Perrier — GaryMarcus · 2026-09-14