Hugging Face researcher: encrypted reasoning is the worst thing to happen to LLM evals

_lewtun · x · 2026-10-06

Lewis Tunstall (lewtun), researcher at Hugging Face, argues that encrypted reasoning is the worst thing to happen to LLM evals, linking to a related discussion.

The claim targets the trend of models inflating performance with long, unreadable hidden chains of thought, which makes it hard for external benchmarks to verify what models are actually doing and whether reported scores are trustworthy. The post itself is brief but flags a significant eval-methodology controversy.

Original post →

More from Models

Models channel →