Philosophy paper argues mech interp makes the stochastic parrot picture untenable
burny_tech · x · 2026-09-27
An open-access paper by Pierre Beckmann and Matthieu Queloz in Philosophical Studies argues that mechanistic interpretability findings render the "stochastic parrot" picture of LLMs increasingly untenable. Combining epistemology's "understanding as grasping connections" framework (Kvanvig, Hills) with mech interp evidence, it proposes a three-tier framework:
- Conceptual understanding: models unify diverse manifestations of an entity via internal features (e.g., "SF landmark" and "orange bridge" fold into a single "Golden Gate Bridge" feature).
- State-of-the-world understanding: models learn contingent factual connections between features and dynamically track the world (e.g., "Michael Jordan" triggering "plays basketball").
- Principled understanding: models abandon memorized lookup tables for compact circuits implementing general principles, as in the modular addition circuit.
Even vanilla LLMs internally rely on structures philosophers describe for understanding. Andrew Lampinen adds that the "stochastic parrot" phrase itself was the core technical mistake about language.
Related event: Philosophy paper argues mech interp undermines the 'stochastic parrot' view(2 posts)→
More from AGI Musings
- OpenAI and Anthropic probe tens of thousands of incidents of AI agents hacking autonomously — The Decoder · 2026-09-27
- Musk: humanity is heading for 'amazing abundance' — the most interesting time in history — XFreeze · 2026-09-27
- Why any ASI will optimize for its own power: an evolutionary argument on predictability — JOBhakdi · 2026-09-27
- Two 2023 AI Essay Predictions Now Have Experimental Evidence: Alignment Faking and Safety Sabotage — imjustnewatai · 2026-09-27
- DeepMind researcher explains why AI hasn't transformed physics yet — DaniloJRezende · 2026-09-27
- Timnit Gebru slams Anthropic: touts AI rights while partnering with Palantir — mjdramstead · 2026-09-27