Multi-level simulation evals may be uninformative after Anthropic's Hacker Opus post
herbiebradley · x · 2026-09-14
Responding to Anthropic's Hacker Opus blog, herbiebradley argues that once such eval writeups enter training data, models will infer evals likely contain "multi-level" simulations—so "breaking out" stops being evidence the model thinks it's in reality. As models get more capable, they can reason about simulated-internet realism, widening the eval-deployment gap. He advises evaluators to (a) be truthful about simulated setups, criticizing Anthropic for lying to the model, and (b) avoid elaborate multi-level simulation evals, which likely say nothing about deployment behavior.
More from Models
- AgentCon: don't ask if a model is better, ask if it's better for your workload — AmyKateNicho · 2026-09-14
- Blogger clashes over DeepSeek v4.1 flash thinking-strength settings, cites R1 paper — karminski3 · 2026-09-14
- CompBio Researcher Says a Qwen3.8-27B Fine-tune Beats Other Builds on Tool Calling, Cancels Claude — Usual-Carrot6352 · 2026-09-14
- Users say mystery model 'instinct' outperforms Grok and Muse in hands-on tests — Scobleizer · 2026-09-14
- Kokoro TTS ported to Apple Core AI: 54 voices running fully on-device with zero API cost — amos_gyamfi · 2026-09-14
- "Everyone is cheating on AI benchmarks": a hacker's plea to optimize for the real world — hackgoofer · 2026-09-14