COLM 2026 paper: recent claims that LLMs can introspect don't meet the evidentiary bar
tallinzen · x · 2026-09-09
A paper accepted at COLM 2026, led by @shashwats19 with @ravfogel and @tallinzen, challenges recent claims about LLM introspection.
- Recent work argued LLMs notice when their "thoughts" are tampered with and can report their contents — i.e., they can introspect;
- The team formalizes the evidentiary burdens implied by commonly accepted definitions of introspection and shows some recently proposed work does not satisfy them;
- Their conclusion: it's too early to claim LLMs can introspect about their internal states.
More from Models
- K2 Horizon open-sources six model scales; 0.9B posts 48.5 on AIME 2026 — kimmonismus · 2026-09-09
- IFM open-sources K2 Horizon: six models, 20T tokens each, and a public reward-hacking audit — kimmonismus · 2026-09-09
- IFM's K2 Horizon: six models from 0.9B to 375B with only 4B/23B active params — kimmonismus · 2026-09-09
- Dev launches aggregator site collecting all statements on OpenAI's claimed Navier–Stokes proof — NathanpmYoung · 2026-09-09
- Meta launches Muse, a personal agent powered by Muse Spark 1.3 that takes actions for you — AIatMeta · 2026-09-09
- GLM-5.3 deep dive: price-performance, AA's new Intelligence Index, Flash architecture — philipkiely · 2026-09-09