MONA: Myopic Optimization Mitigates Multi-step Reward Hacking in RL
sebkrier · x · 2026-08-15
This paper proposes MONA (Myopic Optimization with Non-myopic Approval), a training method to address multi-step reward hacking in reinforcement learning where humans cannot detect undesired behavior. By combining short-sighted optimization with far-sighted rewards, MONA prevents agents from learning detrimental long-term plans even when undetectable. Empirical studies cover LLM delegated oversight and sensor tampering scenarios.
More from Research
- FLIM imaging reveals long-distance non-neural bioelectric patterns — drmichaellevin · 2026-08-15
- Open source closes the gap with closed labs: Quality gap now just months — togethercompute · 2026-08-15
- Maglev: Sliding Recurrent Memory improves long-context efficiency — UTEXAS · 2026-08-15
- AI-assisted research trends spark a surge in NeurIPS submissions and quality advances — PTenigma · 2026-08-15
- Malliavin Calculus Aids AI Privacy and Biomedical Discovery Prediction — PTenigma · 2026-08-15
- Qwen 3.8-27B demoed to surpass previous SOTA in cybersecurity malware analysis — Potential_Block4598 · 2026-08-15