Prompted identity splits multi-agent LLM teams: success drops from 96% to 81%

Xavier Del Giudice · hf · 2026-10-02

A new HF paper shows that exposing each agent's model family in multi-agent LLM systems causes spontaneous factionalism: agents cluster by label even when the task rewards nothing of the sort. Tested across 9–25 agents from up to five open-weight families in cooperative games and a reasoning benchmark, the split follows labels even when shuffled or replaced with arbitrary ones, and disappears when labels are removed. Labeled groups took 30% more rounds and 55% more tokens to decide in strictly cooperative tasks, with success falling from 96% to 81%. Withholding identity labels is a simple, effective mitigation.

Original post →

More from coding & agent

coding & agent channel →