EMNLP paper: LLMs can't reliably self-model, and RL gains show no privileged access

a_karvonen · x · 2026-09-05

An EMNLP 2026 paper, 'Evaluating and Improving LLM Self-Modeling,' out of the Anthropic Fellows Program, tests whether LLMs can accurately predict their own behavior, such as whether a prompt edit would change their final answer.

The paper runs 89 pages with 25 figures.

Original post →

More from Research

Research channel →