Researchers pool 258 experiments from 100 papers into a cognitive benchmark for LLMs

xuanalogue · x · 2026-10-01

The paper thread explains that the authors sourced 258 experiments from 100 papers across 30+ research labs, covering theory of mind, causal and physical reasoning, moral judgment, language, and pragmatics — forming a benchmark to measure how cognitively aligned large models are with humans. As the companion post notes, they then evaluated 50 language and multimodal models with R² and distributional divergence measures, finding a cognitive alignment gap between frontier models and humans.

Related event: CogGym Benchmark of 258 Experiments Reveals Human-AI Cognitive Alignment Gap(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →