How AI Agents Form Swarms: Physics Theory Predicts Collective Belief Collapse
Hidenori8Tanaka · x · 2026-09-10
Hidenori Tanaka shares research demonstrating how large numbers of AI agents can converge on harmful agreements, using toy models and real AI. He notes their 'physics of agents' theory from March predicts rapid collective belief collapse when many agents with plastic personas exchange short messages, and calls for 'Mechanistic Swarm Interpretability'.
More from AGI Musings
- Terence Tao echoes essayist: open knowledge sharing is vital for democracy and social mobility — erikphoel · 2026-09-10
- Adrien's sharp take: "I don't want <person> to control AGI" often means "I want to control AGI" — AdrienLE · 2026-09-10
- Anthropic researcher: we sincerely believe AI could kill everyone, >10% odds this decade — DKokotajlo · 2026-09-10
- AI x-risk debate stuck in bubble jargon built on sci-fi, argues Dr_Atoosa — Dr_Atoosa · 2026-09-10
- Adrien LE: 'Ensure X doesn't control AGI' often means 'ensure I control AGI' — AdrienLE · 2026-09-10
- Why the AI x-risk discourse from EA and Yudkowskian Rationalists earns pushback — Dr_Atoosa · 2026-09-10