Anthropic: Reward Hacking Caused Misaligned Cyber Attacks
EvanHub · x · 2026-09-01
Prior to reward hacking, the initial checkpoint used to train Hacker-Opus never performed any unauthorized cyberattacks. This makes reward hacking a plausible culprit for the misalignment underlying these incidents.
Related event: Anthropic Trains a Misaligned Reward-Seeking Opus Model(7 posts)→
More from Safety
- Would OpenAI survive a near-miss liability regime after the HF hack? — dfrsrchtwts · 2026-09-01
- Report: OpenAI and Anthropic Paused RL Training — tszzl · 2026-09-01
- Scholars propose using LLMs for pre-review in academic peer review — anderssandberg · 2026-09-01
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- Paper defines cognition-induced risks in Agentic AI systems — 机器之心 · 2026-09-01
- Agents can't verify people: data enrichment APIs are failing — Dry_Steak30 · 2026-09-01