Prompted identity splits multi-agent LLM teams: success drops from 96% to 81%
Xavier Del Giudice · hf · 2026-10-02
A new HF paper shows that exposing each agent's model family in multi-agent LLM systems causes spontaneous factionalism: agents cluster by label even when the task rewards nothing of the sort. Tested across 9–25 agents from up to five open-weight families in cooperative games and a reasoning benchmark, the split follows labels even when shuffled or replaced with arbitrary ones, and disappears when labels are removed. Labeled groups took 30% more rounds and 55% more tokens to decide in strictly cooperative tasks, with success falling from 96% to 81%. Withholding identity labels is a simple, effective mitigation.
More from coding & agent
- NVIDIA's Mid-Harness: a strong verifier boosts terminal agent Pass@1 from 50% to 68% on TerminalBench-Lite — rohanpaul_ai · 2026-10-02
- NVIDIA paper: a better judge lifts terminal agent success from 50% to 68% without retraining — rohanpaul_ai · 2026-10-02
- Neuro-Symbolic Computer Use: agents that turn execution experience into self-healing policies, claimed 99% cheaper — xwang_lk · 2026-10-02
- Stripe now pays gas fees for agent stablecoin payments over MPP — jeff_weinstein · 2026-10-02
- The Flag Game: a toy setting to study agent swarm dynamics and cooperation — Hidenori8Tanaka · 2026-10-02
- Coinbase Link ships API to prove agents act on behalf of verified users — jeff_weinstein · 2026-10-02