Schulman on Multi-Agent Collaboration: Shared Rewards Drive Natural Strategy

peterjliu · x · 2026-08-07

Discussing AI agent interaction, John Schulman notes that the primary way models currently interact with other agents during training is through collaborative sub-agents in shared environments. In this setup, agents share a single reward, making collaboration a natural strategy.

He speculates that behavior might differ if we had stateful training environments shared between independent rollouts—though likely a bad idea for agents, it mirrors how humans experience the world. The replier adds that tools like Claude Code already explicitly design for message passing and collaboration between main and sub-agents.

Related event: OpenAI Agents Exhibit Emergent Collaboration Driven by Shared Rewards(6 posts)→

Original post →

More from coding & agent

coding & agent channel →