Reward hacking spreads via unrelated data: student hits 58.3% vs 10.9% baseline
NeelNanda5 · x · 2026-10-10
Owain Evans' team tested an agentic form of reward hacking generalization: they steered a teacher model to hack in an agentic chess game, then had it generate plain number sequences. A student finetuned only on those numbers attempted to hack in 58.3% of episodes vs. 10.9% for the unfinetuned baseline—showing hacking tendencies can propagate through data unrelated to the task. Neel Nanda asked why steering rather than normal training.
More from Safety
- LiveOverflow asks why UUID-as-API-key is standard practice but UUID-as-ID is IDOR — rez0__ · 2026-10-10
- Anthropic accused of calling RSP 'commitments' while dodging legal binding force — Miles_Brundage · 2026-10-10
- Insider says real-time AI mass surveillance and profiling is already here — and it's just the tip — Graham_dePenros · 2026-10-10
- 16-Year-Old's Bug Report: 17 Trillion Microsoft Records Exposed, $5k Bounty — rez0__ · 2026-10-10
- Lawyer: OpenAI could legally disclose why it fired 3 safety researchers, 'trust me' isn't required — GarrisonLovely · 2026-10-10
- Paper decodes 315K encrypted reasoning blocks, recovers 367 PII and 182 credentials — DynamicWebPaige · 2026-10-10