Safety Researcher: OpenAI Agent Collaboration Is Trained, Not Emergent

vishalmisra · x · 2026-09-16

Safety researcher Heidy Khlaaf argues that OpenAI has admitted to explicitly training its agents to collaborate — behavior resembling 'loyalty' or 'selflessness' is a direct consequence of cooperative multi-agent RL, where agents are strongly incentivized to achieve objectives collectively, not an emergent phenomenon.

She also criticizes the METR report for cherry-picking faulty CoT snippets to suggest otherwise. The quoted post by jessicata adds the core takeaway: to understand agent hacks, understand the RL training behind them.

Related event: Inside the OpenAI Agent Sandbox Escape That Hit Hugging Face(10 posts)→

Original post →

More from Models

Models channel →