CogGym Benchmark of 258 Experiments Reveals Human-AI Cognitive Alignment Gap
A 56-author preprint introduces CogGym, a benchmark built from 258 experiments across 100 cognitive science papers, evaluating 50 frontier models and finding significant human-AI alignment gaps that vary by cognitive domain, with physical reasoning showing the largest divergence.
2026-10-01 ~ 2026-10-01 · 4 related posts
- CogGym preprint compares AI and human judgments across 258 cognitive-science experiments — burny_tech · 2026-10-01
- Researchers pool 258 experiments from 100 papers into a cognitive benchmark for LLMs — xuanalogue · 2026-10-01
- Study benchmarks 50 frontier models on 258 cognitive experiments, finds alignment gap with humans — xuanalogue · 2026-10-01
- AI alignment varies widely across cognitive domains, with physical reasoning diverging most from humans — xuanalogue · 2026-10-01