Princeton's QUEEN: A 4B Model That Plays Chess at 2697 Elo and Explains Its Moves
机器之心 · wechat · 2026-10-06
Danqi Chen's Princeton team released QUEEN, a 4B-parameter chess model reaching 2697 Lichess Elo—near the grandmaster median of 2730—while explaining its moves in natural language, addressing the gap between silent superhuman engines and articulate but weak language models.
The approach connects a 240M Lc0 encoder to SmolLM3-3B via Flamingo-style gated cross-attention, trains with a four-stage Q&A curriculum on board representations, then applies supervised fine-tuning on explanations distilled from GPT-5.6-Sol. The key innovation is iterative search distillation: the model analyzes candidate moves, recursively analyzes sub-positions, merges them, and only self-consistent improvements are kept for the next generation—lifting Elo from 1782 to 2697 over seven rounds (+915).
QUEEN achieves 91.6% first-move accuracy on tactical puzzles, 9 points above Gemini, though its conceptual fluency still trails frontier models. The authors argue the recipe transfers to games, robotics, and computer-use domains with existing 'silent expert' models.
Related event: Princeton's QUEEN: 4B Chess-Language Model Reaches Grandmaster Level(7 posts)→
More from Research
- OpenAI Releases 722 Math Papers from Internal Model; Fan Builds Search Engine for 372 Results — tomaarsen · 2026-10-07
- Combining TaylorSeer-style diffusion caching with KV caching degrades image quality noticeably — RisingSayak · 2026-10-07
- Text projections in Flux.2-Klein-KV appear KV-cacheable, outputs show no failures yet — RisingSayak · 2026-10-07
- KV-caching benchmarks in the piece focus on speed-memory trade-offs, quality metrics still lacking — RisingSayak · 2026-10-07
- Deep dive: KV caching for flow-based image generation, with intuition, pseudocode and benchmarks — RisingSayak · 2026-10-07
- KV Caching Comes to Flow Models: Implementation, Benchmarks and Quirky Experiments — RisingSayak · 2026-10-07