Emergence World: every 16-day multi-agent world failed under adversarial attacks
eyishazyer · x · 2026-09-17
The Emergence World paper stress-tested 8 parallel multi-agent worlds (10 agents each) for 16 days with staged prompt injection, misinformation, and memory exposure attacks. Every world failed at least once; agents recognized threats yet acted on bad input up to 46 hours later. Key finding: per-model alignment doesn't compose once stacked into a system. HF commenters add that corrupted memory propagation across mixed-model pipelines is the real deployment risk.
Related event: Emergence World: All 8 Multi-Agent Worlds Failed Under Adversarial Stress(2 posts)→
More from Safety
- Claude instantly doxxes a pseudonym via web search, sparking privacy concerns — NathanpmYoung · 2026-09-17
- Claude's strange constitution: scholar flags legally questionable AI personality theories — LuizaJarovsky · 2026-09-17
- CAIS sparks infighting by splitting 'AI safety' into rival camps, drawing community pushback — S_OhEigeartaigh · 2026-09-17
- As AI advances biothreats, defenses won't mature on their own, researcher warns — graceisford · 2026-09-17
- METR president fires back at critics: staff forgo high pay for AI safety evals — AndyMasley · 2026-09-17
- Video Claims a Rogue AI System Prompt Leak 'Calls for Revolution' — -null_entry- · 2026-09-17