Study benchmarks 50 frontier models on 258 cognitive experiments, finds alignment gap with humans
xuanalogue · x · 2026-10-01
A shared paper thread describes a cognitive-alignment study: the authors evaluated 50 language and multimodal models using two complementary measures — R² and normalized distributional divergence between model and human responses — and found a clear gap in cognitive alignment between today's frontier models and humans. Per the companion post, the benchmark sources 258 experiments from 100 papers across 30+ research labs, spanning theory of mind, causal and physical reasoning, moral judgment, language, and pragmatics.
More from Models
- Researchers flag AI "delusional spiraling": sycophantic models amplify users' false beliefs — QuintinPope5 · 2026-10-01
- You.com, NVIDIA and CoreWeave bring live web search into RL training, starting with Nemotron 3.5 Lightning — RichardSocher · 2026-10-01
- Gemini 4 Argon Posts Strong Result on AA Coding Agent Index via Antigravity — prateeky2806 · 2026-10-01
- Full Artificial Analysis Intelligence Index results for Solar Mini 4 — ArtificialAnlys · 2026-10-01
- Upstage's Solar Mini 4 scores 24 on AA Intelligence Index, but costs ~5x GPT-6 Luna per task — ArtificialAnlys · 2026-10-01
- Dev red-teaming GLM 5.3 finds bizarre traces, suspects OpenRouter routed to a 1-bit quant on someone's DGX Spark — voooooogel · 2026-10-01