Allspark: Alternating Chain-of-Thought Enables Weak-to-Strong Reasoning Transfer
UTEXAS · hf · 2026-09-30
UT Austin researchers introduce Allspark, a training and inference framework for weak-to-strong transfer via alternating chains of thought, motivated by the high cost of large-model RL rollouts.
A weak teacher model is trained alongside a frozen copy of itself; the two alternate reasoning segments, with the frozen model producing final answers. At inference, a stronger student replaces the frozen partner—both remain fixed. Because communication is pure text, the teacher can steer students across model families and tokenizers.
Experiments span controlled Qwen math/reasoning setups and larger Inkling runs on ARC-AGI-2, showing accuracy gains in within- and cross-family transfer (including to Kimi and Nemotron), with varying inference-time tradeoffs between accuracy and tokens.
More from Research
- Diffusion Reward Models: learning the distribution of human preference instead of a single score — burny_tech · 2026-09-30
- Paper rejected over AI-generated review: scholars demand sanctions as trust erodes — paulnovosad · 2026-09-30
- Judea Pearl points to Eric Horvitz's remarks on probabilistic reasoning's 42-year-old rise in AI — yudapearl · 2026-09-30
- Zhuoran Lu lands second HCOMP Honorable Mention, new group is hiring — windx0303 · 2026-09-30
- HCOMP 2026 Honorable Mention: Bayesian cascade analysis of AI credibility indicators — windx0303 · 2026-09-30
- RuneScape Bench Turns an Open-Source MMO Into a Multi-Agent Testbed for LLM Agents — mrdrozdov · 2026-09-30