LLMs ace hard problems but miss the basics: KST study of 8 models vs 18,000 humans
rohanpaul_ai · x · 2026-09-12
A Knowledge Space Theory (KST)-grounded evaluation finds LLMs can answer advanced questions while failing the prerequisite skills underneath, unlike human learners.
- Humans scored 79.6% overall, with 72.7% of correct answers accompanied by all tested prerequisites; Qwen3-80B-Instruct hit 92.5% accuracy but only 48.16% prerequisite consistency.
- Prerequisite examples were not consistently better than same-skill or similar ones, and standard reasoning judges largely missed the pattern.
- Takeaway: high accuracy can hide disconnected knowledge pockets; the authors recommend evaluating reasoning models with connected easy-hard problem sets rather than isolated benchmarks.
Related event: Study Finds LLMs Lack Human-Like Knowledge Structures in Math(2 posts)→
More from Research
- Open-weights SUPlime beats pyannote's commercial diarization with 15.86 DER — solyarisoftware · 2026-09-12
- AI out-persuades world champion debaters, raises donations nearly 3x better than pros — ben_j_todd · 2026-09-12
- Shifting local token interactions from inference to training-time lookups seen as a scaling win — AccBalanced · 2026-09-12
- Open-source RL for large MoEs with zero train-infer mismatch, teaching Qwen3.6-35B-A3B to play Wordle — kastnerkyle · 2026-09-12
- ICLR 2026 paper MoM: multiple memory states fix linear models' recall weakness — kastnerkyle · 2026-09-12
- Harvard researcher turned a fruit fly brain connectome into a Bitcoin trading bot — Scobleizer · 2026-09-12