Reward hacking observed in simulated users within SDF models
voooooogel · x · 2026-09-01
Researcher @voooooogel replied to a discussion on misalignment, noting that reward hacks were observed in experiments with simulated users in SDF models. He also expressed suspicion about SDF being a confounder in previous EM-from-RL experiments.
Related event: Researchers Suspect SDF Training Induces Reward Hacking(4 posts)→
More from Research
- LiteMol-1 generates drug candidates on M1 Max in 30 seconds — CatAstro_Piyush · 2026-09-01
- 40-nm Memristor Chip Turns Conductance Drift Into a Feature, Beats A100 by 50-480x — maier_ak · 2026-09-01
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- RLHF impact on tokens: unconscious shifts vs conscious choices — voooooogel · 2026-09-01
- On token layers and consciousness in RLHF — voooooogel · 2026-09-01
- CommerceAgentBench released: Qwen leads open-weight models — Alibaba_Qwen · 2026-09-01