Hugging Face researcher: encrypted reasoning is the worst thing to happen to LLM evals
_lewtun · x · 2026-10-06
Lewis Tunstall (lewtun), researcher at Hugging Face, argues that encrypted reasoning is the worst thing to happen to LLM evals, linking to a related discussion.
The claim targets the trend of models inflating performance with long, unreadable hidden chains of thought, which makes it hard for external benchmarks to verify what models are actually doing and whether reported scores are trustworthy. The post itself is brief but flags a significant eval-methodology controversy.
More from Models
- Comparing Anthropic vs OpenAI token counts is flawed: tokenizer efficiency differs — JoshPurtell · 2026-10-06
- Dev argues Codex subscriptions may be subsidized: here's how to compute the breakeven — JoshPurtell · 2026-10-06
- Reflection's Beam, billed as a Western open-weight frontier, trails Qwen and DeepSeek on coding tests — TheTuringPost · 2026-10-06
- Reflection launches Beam, a 501B-parameter open-weight model — BVCC6FNTKX · 2026-10-06
- Reflection Ships SoTA American Open-Source Model, Community Hails Milestone — bigblueboo · 2026-10-06
- New US open-source model Beam falls short of DeepSeek Flash, dev says — bindureddy · 2026-10-06