New arXiv paper: current evidence insufficient to show LLMs can introspect

tallinzen · x · 2026-10-02

The paper "Can LLMs Introspect? A Reality Check" (Singh, Linzen, Ravfogel; to be presented at COLM) argues claims of LLM introspection are premature. It proposes two necessary conditions — privileged access (not solvable from input cues) and second-order computation — then re-examines two paradigms: input-only classifiers match models' self-state predictions, and models can't reliably distinguish internal tampering from input manipulation, suggesting generic anomaly detection rather than true introspection. Conclusion: current evidence is insufficient.

Related event: Researchers Challenge Evidence That LLMs Can Introspect(2 posts)→

Original post →

More from Research

Research channel →