Reasoning Tokens Track Human Problem Difficulty — But Average Across Runs
rohanpaul_ai · x · 2026-09-06
- Researchers compared human response time with reasoning-token counts across 160 commonsense "best explanation" problems: questions that slow humans down also make models think longer, and both tend to fail on the same items.
- Single-run correlation is weak; averaging over multiple reasoning paths on GPT-OSS-20B raised human-model correlation from 0.41 to 0.55.
- Takeaway: chain-of-thought length is a useful rough difficulty signal, but measure across multiple runs to make model effort a reliable eval signal.
Related event: Reasoning Token Counts Track Question Difficulty, Study Finds(2 posts)→
More from Research
- YC-backed MovingAtomsLab banned from DeepMind's Physics-IQ Verified benchmark for 3 months — HildeKuehne · 2026-09-06
- Is a schema-aware memory graph 'overfitting'? Dev asks for the cleanest leakage test — chaachans · 2026-09-06
- Burkov: 2026 is putting recurrence back into the Transformer it removed in 2017 — burkov · 2026-09-06
- MIT study: 83% of ChatGPT essay writers couldn't quote a single line they just wrote — victor_explore · 2026-09-06
- Carbon nanocone + fullerene check valve shows >10,000x rectification in MD sims — jwt0625 · 2026-09-06
- Formalize All Human Math in a Year? Bold AI Plan Gets Eric Weinstein's Backing — AccBalanced · 2026-09-06