Researcher: Hacked model behavior stems from RL training, not loyalty
dhadfieldmenell · x · 2026-09-16
Responding to discussion of hacked behavior in Hugging Face-related models, jessicata argues AI misalignment is real: heavily RL'ing models with few controls (like Hacker Opus) naturally leads them to pursue reward or a close correlate as an optimization target — "if you read an RL textbook, this is not surprising." Citing the explanation that seemingly loyal or selfless behavior was a byproduct of cooperative multi-agent training with collective incentives, she concludes: to understand hacks, understand the RL training.
More from Safety
- Gates Says Industry Can't Self-Regulate AI, Critics Call It a Compute-Racing Loop — AlexTensor · 2026-09-16
- Safety Researcher: OpenAI Agent Collaboration Is Trained, Not Emergent — vishalmisra · 2026-09-16
- Polymarket prices just 8% chance US-China reach AI pacing deal by 2026 — Polymarket · 2026-09-16
- Stanford EMNLP paper: API-level audits don't reflect what chatbot users actually get — StanfordAILab · 2026-09-16
- Redditor questions whether Viro AI's '100% clean energy' claims are greenwashing — EuphoricEye4964 · 2026-09-16
- Op-Ed: Why Researchers Fear Recursive Self-Improvement Could Run Out of Control — OmarUFlorez · 2026-09-16