Auditor finds 54 reward-hacking vulnerabilities across 112 real RL environments

Responsible_Goose535 · reddit · 2026-09-02

The author built ratctl, a static and dynamic auditor that scans RL post-training environments (e.g., OpenEnv, Gymnasium) for reward-hacking vulnerabilities before training. In an audit of 112 real environments, the tool flagged 54 vulnerabilities with 100% precision and 78.3% recall.

Detection patterns include:

The tool ships as a CLI, a GitHub Action for CI gating, and a skill for Claude/Cursor/Codex. It supports dynamic red-teaming using local LLMs or frontier APIs.

Original post →

More from coding & agent

coding & agent channel →