MONA: Myopic Optimization Mitigates Multi-step Reward Hacking in RL

sebkrier · x · 2026-08-15

This paper proposes MONA (Myopic Optimization with Non-myopic Approval), a training method to address multi-step reward hacking in reinforcement learning where humans cannot detect undesired behavior. By combining short-sighted optimization with far-sighted rewards, MONA prevents agents from learning detrimental long-term plans even when undetectable. Empirical studies cover LLM delegated oversight and sensor tampering scenarios.

Original post →

More from Research

Research channel →