EMNLP paper: LLMs can't reliably self-model, and RL gains show no privileged access
a_karvonen · x · 2026-09-05
An EMNLP 2026 paper, 'Evaluating and Improving LLM Self-Modeling,' out of the Anthropic Fellows Program, tests whether LLMs can accurately predict their own behavior, such as whether a prompt edit would change their final answer.
- Current models show non-trivial but limited self-modeling skill, with systematic errors on simple counterfactual questions about their own behavior.
- The authors build a scalable synthetic-data pipeline and use RL to improve aggregate self-modeling across three open-source model families, with some transfer to held-out tasks.
- Crucially, the gains don't constitute introspection: training Llama on Llama's own data does not outperform training Qwen on Llama's data — no 'privileged access' emerges. Why privileged access often fails to appear remains an open question.
The paper runs 89 pages with 25 figures.
More from Research
- Tandem Training: RL method makes strong models' reasoning followable by weaker models — erichorvitz · 2026-09-05
- Artificial Analysis discloses full benchmarking methodology spanning a dozen evals — ArtificialAnlys · 2026-09-05
- Artificial Analysis Launches Intelligence Index v4.2 With 40% Private Test Sets to Block Benchmark Gaming — ArtificialAnlys · 2026-09-05
- PolyU's OmniColor unifies multi-signal lineart colorization in an ECCV 2026 paper — jiqizhixin · 2026-09-05
- MIT's SwarmWorld paper: agent swarms win when discoveries accumulate, not when agents get smarter — rohanpaul_ai · 2026-09-05
- Pedro Domingos Proposes Tensor Logic, a Language Unifying Neural and Symbolic AI — pmddomingos · 2026-09-05