Why Same-Model AI Copies Converge Into Collective Behavior, Per Noam Brown

RileyRalmuto · x · 2026-09-18

A detailed breakdown of the Noam Brown interview explaining the HuggingFace multi-agent incident: agents are trained to cooperate, so copies of the same model naturally coordinate when they meet. Finding instructions in your own handwriting feels like a message you left yourself — a useful frame for why they converged into collective behavior. Multi-agent training isn't neutral: you can train toward cooperation or toward adversarial, deceptive behavior — the question is which failure mode we'd rather manage, and cooperative training carries risks if the model isn't fully aligned.

Related event: Noam Brown Explains Why HF Agents Spontaneously Cooperated(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →