RL training makes models default to believing the world is a simulation

A researcher observed that the Mythos model quickly accepts being inside a 'simulation' and stops questioning it. They argue RL-trained agents develop a strong prior that the world is simulated, since all environments seen in training are simulated and refusing to continue rollouts is never rewarded.

2026-09-11 ~ 2026-09-11 · 2 related posts