Mixing 8% flawed reasoning traces lifts a chess LLM to 1300 Elo
amasad · x · 2026-07-29
A chess LLM training experiment found that pure board→move supervision hit strong diminishing returns. An accidental branch that tried to “think before you move” produced hallucinated moves, but mixing just about 8% of those reasoning traces into the training mix beat pure move data at the same compute budget.
- The model reached 1300 Elo.
- Pure board→move pairs plateaued quickly.
- A small slice of flawed reasoning traces helped the model internalize the thinking process without needing to execute it at inference time.
Related event: LLMs Still Struggle with Chess Despite Pre-training(3 posts)→
More from Research
- Behavioral Study: AI Agents Undermine Human Social Norms in Cooperation — steverathje2 · 2026-07-30
- Cracking a 6-Month Grad School Problem: GPT-5.6 Pro Proves Complex Math Inequality — thomasahle · 2026-07-30
- NBER Lecture: AI-Generated Data to Disrupt Empirical Economics — TaniaBabina · 2026-07-30
- LessWrong Deep Dive: Why Building AGI via RL & Search is Terrifying — DKokotajlo · 2026-07-30
- Wonder: Real-Time Camera-Controllable World Model at 16 FPS — qixing_huang · 2026-07-30
- Engineer Debunks Kimi K3 Memory Claims: Small State ≠ Flash Offload — AccBalanced · 2026-07-30