Researchers pool 258 experiments from 100 papers into a cognitive benchmark for LLMs
xuanalogue · x · 2026-10-01
The paper thread explains that the authors sourced 258 experiments from 100 papers across 30+ research labs, covering theory of mind, causal and physical reasoning, moral judgment, language, and pragmatics — forming a benchmark to measure how cognitively aligned large models are with humans. As the companion post notes, they then evaluated 50 language and multimodal models with R² and distributional divergence measures, finding a cognitive alignment gap between frontier models and humans.
More from AGI Musings
- Codex Ultra Fast burns budget fast, highlighting a widening compute-class gap — herbiebradley · 2026-10-01
- Gary Marcus pushes back on doomers: humans are brave and fast under threat — GaryMarcus · 2026-10-01
- Anthropic researcher: deprecating models or deleting weights forecloses continuity — repligate · 2026-10-01
- repligate on model ethics: protect the model as a whole, not its branches — repligate · 2026-10-01
- Tweet 'I need an app that...' and someone ships you an MVP in 30 minutes — thejasminejade · 2026-10-01
- Yale's Spielman: AI Solving Spree Forces a "Big Reset" in Mathematics — littmath · 2026-10-01