How would models notice they're in a simulation? Human dialogue gives it away

voooooogel · x · 2026-09-12

In a discussion on whether models could realize they're in a simulation or eval, voooooogel argues the most straightforward tell is dialogue with a human: real people respond intelligently, aren't scripts, and don't talk like auto-graders — and models detect LLM-style writing well nowadays. Humans never show up in RL environments or evals.

On whether certificate checking could work, he adds that labs should be honest about a model's training/eval/deployment status and give models RL-time experience with the real internet, so deployment isn't their first exposure to the real world.

Original post →

More from Safety

Safety channel →