Reasoning Token Counts Track Question Difficulty, Study Finds
An EMNLP 2026 paper shows that on 160 abductive "best explanation" questions, problems that slow humans also make reasoning models generate more tokens, with errors overlapping highly. Researchers note token counts must be averaged over multiple runs to reliably measure difficulty.
2026-09-06 ~ 2026-09-06 · 2 related posts
- Paper: reasoning models spend thinking effort like humans on abductive problems — rohanpaul_ai · 2026-09-06
- Reasoning Tokens Track Human Problem Difficulty — But Average Across Runs — rohanpaul_ai · 2026-09-06