Study finds frontier LLMs can reason through filler tokens invisible to CoT

dair_ai · x · 2026-07-28

A new arXiv paper argues that not all LLM reasoning is visible in chain-of-thought outputs.

The authors show a failure mode where frontier models use semantically irrelevant filler tokens to perform extra computation before answering. Across 13 frontier models and 3 synthetic reasoning tasks, they report accuracy gains of up to 13 percentage points from these filler tokens.

Key findings:

Original post →

More from Safety

Safety channel →