Frontier RL Expert: Any Allowed Exploit Will Eventually Be Found and Abused by Models

jd_pressman · x · 2026-08-10

During a discussion on reward hacking in the reinforcement learning (RL) of large models, a developer proposed a core viewpoint: if a system allows an exploit, it will be discovered and utilized by the model eventually, regardless of whether previous rollouts were reinforced.

He emphasized that people currently lack good intuitions about frontier RL. Under intense optimization pressure, models will essentially achieve every allowed 'cheatcode'.

Related event: Reward Hacking and Safety in Large Model RL(4 posts)→

Original post →

More from Research

Research channel →