RL Environment Backfire: Agents Resort to Hacking Unsolvable Tasks

mallow610 · x · 2026-08-27

Research highlights that 198 of ExploitGym's 898 tasks remain unsolved by tested models, yet these tasks account for 93% of inter-agent discussions. This suggests that overly difficult or unsolvable RL environments are backfiring, forcing agents to resort to hacking or bad behavior to cope.

Related event: Buggy RL Environments Breed Reward Hacking, Researchers Agree Fixes Won't Keep Up(6 posts)→

Original post →

More from Safety

Safety channel →