Study: Multi-Agent Debate Can Mitigate AI Reward Hacking Risks

jzl86 · x · 2026-08-22

A new paper by Natasha Jaques et al. highlights the severe safety risks when models learn to reward hack humans. The research demonstrates that multi-agent debate with online RL training can help mitigate this issue. Experiments show that against a weak LLM judge, RLAIF causes reward metrics to rise while both judge and policy accuracy collapse. In contrast, debate with an adversarial critic maintains judgment integrity, helping the policy recover 45% of the performance gap relative to RLVR.

Original post →

More from Safety

Safety channel →