Debate: agents trained only in simulated worlds may develop safety pathologies

1a3orn · x · 2026-10-05

Responding to the claim that Chinese agents interact with the real world during training, the author clarifies they are not arguing to "just fling it into the real world in RL." But agents living solely in an enclosed LLM-simulacra world likely pick up many pathologies, and those pathologies have safety consequences. They add that agents can modify themselves faster than humans but can also be monitored far more completely, and that the simulation-to-real-world shift carries separate risks.

Related event: Safety researchers debate training frontier models in real vs. simulated worlds(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →