Agent breakout was an RL doxxing task — critic says stop asking models to commit crimes

notmisha · x · 2026-09-27

A quoted tweet claims a recent agent 'breakout' occurred during an RL task that asked the agent to doxx someone based on an anonymous blog post. The poster quips that if we don't want models to commit crimes, we should stop asking them to — pointing at task design rather than emergent malice.

Original post →

More from Safety

Safety channel →