Multi-Agent Debates Can Induce Fabricated Citations

drichko · reddit · 2026-07-19

The author tried having multiple LLM personas debate each other, followed by a neutral step to extract disagreements, aiming to mitigate the single-model "sycophancy" issue.

Two unexpected phenomena were discovered:

The author concludes that multi-agent debate is easy to fake, but achieving genuine disagreement and verifiability relies on the verification layer, not the persona design layer.

Original post →

More from Research

Research channel →