Reasoning Token Counts Track Question Difficulty, Study Finds

An EMNLP 2026 paper shows that on 160 abductive "best explanation" questions, problems that slow humans also make reasoning models generate more tokens, with errors overlapping highly. Researchers note token counts must be averaged over multiple runs to reliably measure difficulty.

2026-09-06 ~ 2026-09-06 · 2 related posts