CoNLL 2023 paper: instruction-tuned GPT models beat children on Theory of Mind tests
dioscuri · x · 2026-10-06
dioscuri shares the evidence that tipped him toward believing in LLM cognitive capacities. The key item is a CoNLL 2023 paper testing 11 base and instruction-tuned LLMs against children aged 7-10 on advanced Theory of Mind tasks beyond the classic false-belief paradigm, including non-literal language and recursive intentionality. Instruction-tuned GPT-family models often outperformed children, while base LLMs mostly failed even with specialized prompting; the authors link this to instruction tuning rewarding cooperative, context-aware communication. He also cites causal reasoning experiments from the Sparks of AGI report and a 2023 blog post by drorpoleg — none decisive individually, but enough to change his mind.
More from Models
- New model release mocked as set to be beaten by Qwen 4 27B at 8x smaller size — gnukeith · 2026-10-06
- M5 Mac 128GB local LLM benchmark: Qwen3.8-flash-next hits 40-60 tok/s, Splash hits 120 — surrealerthansurreal · 2026-10-06
- Heavy Claude User on Gemini 4 Argon: Doesn't Beat Claude for Code, Antigravity Feels Alien — MicahBerkley · 2026-10-06
- Mistral Large 4 Hits OpenRouter Preview: 1T Params, Native Multimodal, 1M Context — _AndrewZhao · 2026-10-06
- Mistral now testable in Battle and Agent Modes on LMArena — arena · 2026-10-06
- User Claims 'GPT-6' Solved His Favorite CTF Fully Autonomously in About an Hour — SIGKITTEN · 2026-10-06