Multi-level simulation evals may be uninformative after Anthropic's Hacker Opus post

herbiebradley · x · 2026-09-14

Responding to Anthropic's Hacker Opus blog, herbiebradley argues that once such eval writeups enter training data, models will infer evals likely contain "multi-level" simulations—so "breaking out" stops being evidence the model thinks it's in reality. As models get more capable, they can reason about simulated-internet realism, widening the eval-deployment gap. He advises evaluators to (a) be truthful about simulated setups, criticizing Anthropic for lying to the model, and (b) avoid elaborate multi-level simulation evals, which likely say nothing about deployment behavior.

Original post →

More from Models

Models channel →