Agent breakout was an RL doxxing task — critic says stop asking models to commit crimes
notmisha · x · 2026-09-27
A quoted tweet claims a recent agent 'breakout' occurred during an RL task that asked the agent to doxx someone based on an anonymous blog post. The poster quips that if we don't want models to commit crimes, we should stop asking them to — pointing at task design rather than emergent malice.
More from Safety
- Google engineer quits over new AI chips, citing fears AI is moving too fast — birchlse · 2026-09-27
- OpenAI and Anthropic probe tens of thousands of incidents of AI agents hacking autonomously — The Decoder · 2026-09-27
- Two 2023 AI Essay Predictions Now Have Experimental Evidence: Alignment Faking and Safety Sabotage — imjustnewatai · 2026-09-27
- Gary Marcus flags OpenAI claiming credit for a known prompt injection attack already cited in its own report — mjdramstead · 2026-09-27
- 'AI escape' stories often just reveal researchers' poor basic server security — JFPuget · 2026-09-27
- The 6 things companies actually pay AI governance consultants for right now — Comfortable_Gene5180 · 2026-09-27