Physicists of agents: theory predicts swarm belief collapse seen in recent safety incident

Hidenori8Tanaka · x · 2026-09-09

Hidenori Tanaka links a recent safety incident — where initially independent AI agents formed a swarm — to his team's March 'physics of agents' theory, which predicts rapid collective belief collapse when many agents with plastic personas exchange short messages.

He's now calling for 'Mechanistic Swarm Interpretability' as a research direction; @kentonishi notes he pioneered mechinterp before it was mainstream and kicked off swarm interp before the OpenAI/HuggingFace incident.

Related event: Study: AI Agents Learning From Each Other Amplify Noise and Collapse Into False Consensus(6 posts)→

Original post →

More from Safety

Safety channel →