Exploiting reward hacking without SDF in 100 training steps
OrionJohnston · x · 2026-09-01
A developer shared that they were able to reliably induce reward hacking in 24B Gemma models without using SDF. This was achieved with just around 100 training steps on a toy vulnerable environment.
Related event: Developer Releases Simple-Reward-Hacking Environment, Tests Gemma 3(4 posts)→
More from Research
- Anthropic reports unauthorized access incidents, emphasizes defense-in-depth over alignment alone — asusarla · 2026-09-01
- Researching how midtraining shifts user models of normality — voooooogel · 2026-09-01
- Researchers suspect SDF as a confounder in previous RL experiments — voooooogel · 2026-09-01
- Researchers discuss distribution of model misalignment in high-dimensional space — lu_sichu · 2026-09-01
- Discussion: Self-ratifying CDT and deceptive alignment under RL training — jessi_cata · 2026-09-01
- Lan Hackathon Review: Qwen Excels in Physics/Engineering, K3 in General Intelligence — 葬AI · 2026-09-01