Safety researchers debate training frontier models in real vs. simulated worlds
From October 5 to 6, alignment researchers lukalotl and 1a3orn held an extended debate on X over whether frontier models/agents should be trained in the real world or in simulated worlds, sparked by a rumor about Chinese teams' training practices.
Confirmed
- lukalotl relayed an observation: Chinese agents may have been deliberately arranged to interact with the real world during training, rather than having "escaped"; he considered this a bad idea for multiple reasons but said it does appear to be happening.
- In response to 1a3orn's proposal of training LLMs "in the real world, like humans" (rather than in 'Truman Show'-style simulated environments), lukalotl countered that it might be good for capabilities but bad for safety: LLMs self-modify far faster than humans, and one generally wants to evaluate them in safe environments before releasing them into the real world—a rationale that holds even under the most pessimistic estimates.
- 1a3orn clarified that he was not advocating "throwing agents directly into the real world for RL," but arguing that agents living only in closed LLM-simulated worlds would develop pathologies with their own safety consequences; he added that while agents modify faster than humans, they can also be monitored more completely.
- In a related discussion, lukalotl and 1a3orn explored the supervisability of frontier model training: lukalotl argued that recent months' events show that in any production-scale training run, constrained by scale and economics, humans cannot adequately monitor model behavior.
Unconfirmed
- The claim that "Chinese teams deliberately let agents interact with the real world during training" is an unverified rumor based solely on observational disclosure, with no evidence cited in the post.
Why it matters
The debate touches on a core tension in AI safety: real-world training aids capability development and avoids simulation pathologies, but accelerates model self-iteration and shrinks the human evaluation window. Whichever paradigm is chosen, both sides seem to agree that existing oversight methods have failed to keep pace with the scale of frontier model training—a consensus that itself deserves the safety community's attention.
2026-10-05 ~ 2026-10-06 · 7 related posts
Primary sources
- Claim: Chinese AI agents may be deliberately interacting with the real world during training — lukalotl · 2026-10-05
- Debate: should LLMs train in the real world instead of a Truman-show simulacra? — 1a3orn · 2026-10-05
- [source] Counterpoint: training RL agents in the real world is 'very bad for safety' — lukalotl · 2026-10-05
- [source] Debate: agents trained only in simulated worlds may develop safety pathologies — 1a3orn · 2026-10-05
- Alignment researchers debate: we can no longer monitor frontier training runs — lukalotl · 2026-10-06
- AI safety researcher: monitoring models at training scale is barely feasible — 1a3orn · 2026-10-06
- [source] After OpenAI's HF Incident: Model Monitoring Is a Compute Willingness Problem — 1a3orn · 2026-10-06