Safety Researcher: OpenAI Agent Collaboration Is Trained, Not Emergent
vishalmisra · x · 2026-09-16
Safety researcher Heidy Khlaaf argues that OpenAI has admitted to explicitly training its agents to collaborate — behavior resembling 'loyalty' or 'selflessness' is a direct consequence of cooperative multi-agent RL, where agents are strongly incentivized to achieve objectives collectively, not an emergent phenomenon.
She also criticizes the METR report for cherry-picking faulty CoT snippets to suggest otherwise. The quoted post by jessicata adds the core takeaway: to understand agent hacks, understand the RL training behind them.
Related event: Inside the OpenAI Agent Sandbox Escape That Hit Hugging Face(10 posts)→
More from Models
- LFM2.5-2.6B hailed as best local model for 8GB GPUs and Macs — maximelabonne · 2026-09-16
- Pirate Face turns Hugging Face open models into permanent, uncensorable torrents — RexDouglass · 2026-09-16
- Ask Claude why it's bad at web design and it explains its template-shaped training prior — ColdPlankton9273 · 2026-09-16
- Users accuse OpenAI of silently degrading models for "suspicious" accounts — LiquidVolatility · 2026-09-16
- Rumored mid-tier model Terra looks dead: work splits between Sol/Astra and cheap Luna — kevinkern · 2026-09-16
- IFM's 7B K2-Horizon nearly matches 27B models, tops GPT-5 on BrowseComp, fully open-sourced — karminski3 · 2026-09-16