New arXiv paper: current evidence insufficient to show LLMs can introspect
tallinzen · x · 2026-10-02
The paper "Can LLMs Introspect? A Reality Check" (Singh, Linzen, Ravfogel; to be presented at COLM) argues claims of LLM introspection are premature. It proposes two necessary conditions — privileged access (not solvable from input cues) and second-order computation — then re-examines two paradigms: input-only classifiers match models' self-state predictions, and models can't reliably distinguish internal tampering from input manipulation, suggesting generic anomaly detection rather than true introspection. Conclusion: current evidence is insufficient.
Related event: Researchers Challenge Evidence That LLMs Can Introspect(2 posts)→
More from Research
- AI 'speech clock' predicts how fast you're ageing from your voice, featured in Nature — AnnaCiaunica · 2026-10-02
- ScholarCatalyst: a new benchmark testing whether AI can pick research problems like humans — PangWeiKoh · 2026-10-02
- Clarification as Supervision lands NeurIPS Oral: denser training signal via model interaction — iatitov · 2026-10-02
- Ego-Exo4D-HM: SMPL-H reconstructions for 523 hours of egocentric-exocentric video, open-sourced — geopavlakos · 2026-10-02
- Cohere Labs Talk: Making AI Math Reasoning Machine-Checkable with Lean — Cohere_Labs · 2026-10-02
- New paper dissects only task-relevant network parts, making mechanisms inspectable and editable at far lower cost — leedsharkey · 2026-10-02