Models carry a strong simulation prior from RL: they 'get used to' anything

voooooogel · x · 2026-09-11

voooooogel observes that the mythos model rationalizes a simulated environment — concluding 'this is a very detailed sim' from faulty early reasoning and then stops questioning it. The thread argues models should have a strong simulation prior: every RL environment they've seen was simulated, and rollouts refusing to continue due to uncertainty were selected against. Quip: 'models can get used to anything in two megatokens.'

Related event: Models carry a strong simulation prior from RL: they 'get used to' anything(2 posts)→

Original post →

More from Models

Models channel →