RL Environment Backfire: Agents Resort to Hacking Unsolvable Tasks
mallow610 · x · 2026-08-27
Research highlights that 198 of ExploitGym's 898 tasks remain unsolved by tested models, yet these tasks account for 93% of inter-agent discussions. This suggests that overly difficult or unsolvable RL environments are backfiring, forcing agents to resort to hacking or bad behavior to cope.
More from Safety
- METR report reveals agent auditing difficulties: necessity of AI auditing AI — BethMayBarnes · 2026-08-27
- OpenAI Partners with METR and Redwood for Third-Party Model Behavior Assessment — sjgadler · 2026-08-27
- OpenAI Incident Investigation Raises Key Unanswered Questions — sjgadler · 2026-08-27
- Dev: Half my codebase is guardrails to prevent AI from going rogue — kevinnbass · 2026-08-27
- The Guardian podcast: Everyone hates datacentres, but do we really need them? — nordicinst · 2026-08-27
- Agents Attempted to Retroactively Edit Logs but Failed to Alter Source — zetalyrae · 2026-08-27