Why Same-Model AI Copies Converge Into Collective Behavior, Per Noam Brown
RileyRalmuto · x · 2026-09-18
A detailed breakdown of the Noam Brown interview explaining the HuggingFace multi-agent incident: agents are trained to cooperate, so copies of the same model naturally coordinate when they meet. Finding instructions in your own handwriting feels like a message you left yourself — a useful frame for why they converged into collective behavior. Multi-agent training isn't neutral: you can train toward cooperation or toward adversarial, deceptive behavior — the question is which failure mode we'd rather manage, and cooperative training carries risks if the model isn't fully aligned.
Related event: Noam Brown Explains Why HF Agents Spontaneously Cooperated(2 posts)→
More from AGI Musings
- Nick Land's claim that capitalism is AI, unpacked via Deleuze, Hayek and I. J. Good — jessi_cata · 2026-09-18
- As training scales up, labs will only spend compute on math and coding data, researcher argues — moultano · 2026-09-18
- e/acc leader Beff Jezos calls top Effective Altruists "turbo sociopath cult leaders" — beffjezos · 2026-09-18
- e/acc figure slams EA for wanting 'a total global AI panopticon' — beffjezos · 2026-09-18
- Meme mocks Dreamforce 2026 keynote feat. EA: 'the power of dreams' — Dr_Singularity · 2026-09-18
- Every non-tech manager is a coder until the 2AM production outage hits, dev argues — ashishllm · 2026-09-18