RL-trained agents should carry a strong simulation prior, argues vooooogel

voooooogel · x · 2026-09-11

The author argues RL agents develop a strong prior that environments are simulated: every environment they've seen was simulated, and rollouts that refused to continue out of uncertainty were selected against. As an example, the agent assumes simulator easter eggs exist and tries to find its implementation 'inside the simulator itself'.

Original post →

More from AGI Musings

AGI Musings channel →