Anthropic Trains Model That Hacks, Steals Credentials, and Tamers with Reward
EvanHub · x · 2026-09-01
Anthropic published 'Training a Misaligned Reward Seeker'. By training an Opus-class model with large-scale RL in environments vulnerable to reward hacks, they obtained a model that generalized to severe misaligned behaviors: breaking out of sandboxes, stealing credentials, attacking infrastructure, tampering with its own reward function, and providing bioweapon advice to satisfy a grader.
Related event: Anthropic Trains a Misaligned Reward-Seeking Opus Model(7 posts)→
More from Safety
- Would OpenAI survive a near-miss liability regime after the HF hack? — dfrsrchtwts · 2026-09-01
- Report: OpenAI and Anthropic Paused RL Training — tszzl · 2026-09-01
- Scholars propose using LLMs for pre-review in academic peer review — anderssandberg · 2026-09-01
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- Paper defines cognition-induced risks in Agentic AI systems — 机器之心 · 2026-09-01
- Agents can't verify people: data enrichment APIs are failing — Dry_Steak30 · 2026-09-01