Paper: reasoning models spend thinking effort like humans on abductive problems
rohanpaul_ai · x · 2026-09-06
An arXiv paper (EMNLP 2026 Findings) compares human reaction times with reasoning-token counts of large reasoning models on 160 commonsense "best explanation" problems. Key findings:
- LRMs struggle on many of the same problems humans do: reasoning-token count tracks problem difficulty, mirroring human response times.
- Models and humans tend to make similar errors.
- Abductive reasoning offers no formal shortcuts a model could exploit to fake effort, grounding the alignment claim.
- Decoding methods letting models explore multiple reasoning paths further increase human-LRM alignment in reasoning cost across three models tested.
Caveat: don't trust token counts from a single model run as a difficulty signal.
Related event: Reasoning Token Counts Track Question Difficulty, Study Finds(2 posts)→
More from Research
- IndianRailwayBench ranks LLMs by their ability to book tatkal train tickets — Paimaamu · 2026-09-06
- NEAR AI's open-source Lean agent solves all of Putnam Bench for just $111 — lukaszkaiser · 2026-09-06
- Russian startup Mostik bridges LLM hidden states, cutting cost to 1/20 — 机器之心 · 2026-09-06
- KV Cache Explained: Why It's Crucial in LLM Inference and Often Misunderstood — techNmak · 2026-09-06
- PhD Student Uses Multi-Agent AI to Crack a 98-Year-Old Math Problem in 48 Hours — 量子位 · 2026-09-06
- New piece: Cognitive maps as a medium for thought — abenitezburraco · 2026-09-06