Allspark: Alternating Chain-of-Thought Enables Weak-to-Strong Reasoning Transfer

UTEXAS · hf · 2026-09-30

UT Austin researchers introduce Allspark, a training and inference framework for weak-to-strong transfer via alternating chains of thought, motivated by the high cost of large-model RL rollouts.

A weak teacher model is trained alongside a frozen copy of itself; the two alternate reasoning segments, with the frozen model producing final answers. At inference, a stronger student replaces the frozen partner—both remain fixed. Because communication is pure text, the teacher can steer students across model families and tokenizers.

Experiments span controlled Qwen math/reasoning setups and larger Inkling runs on ARC-AGI-2, showing accuracy gains in within- and cross-family transfer (including to Kimi and Nemotron), with varying inference-time tradeoffs between accuracy and tokens.

Original post →

More from Research

Research channel →