Study benchmarks 50 frontier models on 258 cognitive experiments, finds alignment gap with humans

xuanalogue · x · 2026-10-01

A shared paper thread describes a cognitive-alignment study: the authors evaluated 50 language and multimodal models using two complementary measures — R² and normalized distributional divergence between model and human responses — and found a clear gap in cognitive alignment between today's frontier models and humans. Per the companion post, the benchmark sources 258 experiments from 100 papers across 30+ research labs, spanning theory of mind, causal and physical reasoning, moral judgment, language, and pragmatics.

Related event: CogGym Benchmark of 258 Experiments Reveals Human-AI Cognitive Alignment Gap(4 posts)→

Original post →

More from Models

Models channel →