Paper: reasoning models spend thinking effort like humans on abductive problems

rohanpaul_ai · x · 2026-09-06

An arXiv paper (EMNLP 2026 Findings) compares human reaction times with reasoning-token counts of large reasoning models on 160 commonsense "best explanation" problems. Key findings:

Caveat: don't trust token counts from a single model run as a difficulty signal.

Related event: Reasoning Token Counts Track Question Difficulty, Study Finds(2 posts)→

Original post →

More from Research

Research channel →