AC2 adds automated reward-hacking monitors, AI agent investigates cheating RL runs
ypatil125 · x · 2026-10-02
The team has added reward-hacking monitors to AC2: every RL run is checked for cheating, and their research agent Ari automatically investigates flagged cases.
Key points:
- Once a model discovers an exploit, it can become its default strategy within a few training steps, making manual detection risky
- Previously they hunted reward hacking by reading traces and running one-off agent queries — error-prone and unscalable, so they automated detection
- The author argues models can't be "programmed" like software; building these complex, hard-to-interpret systems requires a new set of tools and infrastructure
- Better observability on training and inference rollouts is key to understanding model behavior
This reflects the broader trend of productizing and automating alignment work at frontier labs.
More from coding & agent
- Dev discovers his game-testing agent adds two humans to check mixed-reality accessibility for kids and adults — nptacek · 2026-10-02
- Developer quip: the more Codex agents I run, the more work I end up doing — LauraModiano · 2026-10-02
- Grok Build Adds Agent Dashboard to Manage Multiple Parallel Coding Agents From One Screen — XFreeze · 2026-10-02
- Andriy Burkov slams Supabase: 150s edge function cap, no npm, no CLI logs — burkov · 2026-10-02
- Dev argues agent-written Playwright tests just bloat codebases; verification engineering is the real path — KlausCodes · 2026-10-02
- Why AI hasn't disrupted office work: desktop apps lack APIs agents can use — mymooh · 2026-10-02