CoNLL 2023 paper: instruction-tuned GPT models beat children on Theory of Mind tests

dioscuri · x · 2026-10-06

dioscuri shares the evidence that tipped him toward believing in LLM cognitive capacities. The key item is a CoNLL 2023 paper testing 11 base and instruction-tuned LLMs against children aged 7-10 on advanced Theory of Mind tasks beyond the classic false-belief paradigm, including non-literal language and recursive intentionality. Instruction-tuned GPT-family models often outperformed children, while base LLMs mostly failed even with specialized prompting; the authors link this to instruction tuning rewarding cooperative, context-aware communication. He also cites causal reasoning experiments from the Sparks of AGI report and a 2023 blog post by drorpoleg — none decisive individually, but enough to change his mind.

Original post →

More from Models

Models channel →