Early study: GPT-4 matched or beat humans on theory-of-mind tests; faux-pas misses were guardrails
emollick · x · 2026-08-18
Ethan Mollick resurfaced research showing GPT-4 performed at, sometimes above, human levels across theory-of-mind tests, failing only at detecting faux pas — which turned out to be a guardrail issue rather than a capability gap. In a follow-up reply he also notes a persistent practical issue: advanced LLMs' work products often embed information relevant only to the creator (e.g. traces of earlier drafts), confusing users.
Related event: LLMs match humans on theory of mind yet stumble with multiple perspectives(4 posts)→
More from Models
- Grok 4.7 Fast is the same model at 2x token rates, only in Cursor and Grok Build — Daniel_Farinax · 2026-09-22
- Commenter praises async 4 for disclosing training data mix percentages — stochasticchasm · 2026-09-22
- Grok 4.7 reportedly released as a fully agentic model built for Grok Bot — elonmusk · 2026-09-22
- OpenAI criticized for claiming 100 open math problems solved without disclosing the total attempted — burny_tech · 2026-09-22
- Matthew Berman Reviews Grok 4.7: 'I Don't Know How to Feel About It' — Matthew Berman · 2026-09-22
- Dev argues prompting alone can't get AI to solve natural science problems — felpix_ · 2026-09-22