New paper: LLMs can get answers right while their chain-of-thought traces are invalid
rao2z · x · 2026-10-01
A new paper, "Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces", tests whether a model's reasoning is actually right when its answer is. Using synthetic grade-school math where every trace step is programmatically verifiable, the authors show LLM think traces lack end-user interpretable semantics even on iGSM — a benchmark specifically built by Allen-Zhu to showcase intermediate token semantics. The finding extends their ICML/ACL/TMLR line of work and challenges assumptions that CoT traces are faithful explanations.
More from Research
- Lior and Stanford launch $100K MMBU Challenge to test if biomedical AI really sees — LiorOnAI · 2026-10-01
- Smart contract security audit dataset with 10K-100K findings trends on Hugging Face — Zaevlad · 2026-10-01
- FRAC replaces exponential forgetting with power-law memory in state space models — Ivan Kobyzev · 2026-10-01
- HCOMP 2026 paper: error decomposition for human-in-the-loop evaluation of LLM data annotation — windx0303 · 2026-10-01
- Research shows RL training breaks defenses against distillation attacks, evals give false security — terryyuezhuo · 2026-10-01
- True Positive Weekly #180: Xiaomi's MIT-licensed MiMo-V2.6, physicist-style LLM pruning, OpenHands — burkov · 2026-10-01