Allspark's Weak Alternating CoT Lifts Frozen Kimi by 11.5 Points, No Strong Rollouts

Allspark: Weak to Strong Transfer via Alternating Chain of Thought

Kaizhao Liang, Junxiong Wang, Chen Liang, Zhendong Wang, Qiang Liu

cs.LG, cs.AI

2026-09-27

Allspark trains a weak teacher via alternating CoT, then steers a frozen strong student. Inkling on ARC-AGI-2 peaks at 81.6% vs 78.1% alone; the same teacher lifts Kimi by 11.5 pp.

What problem this solves

Screening an RL recipe on a frontier model is expensive before it is informative. The bottleneck is rollouts, especially long reasoning traces. Labs that cannot pay for that loop cannot even try.

The usual workaround grows the skill in a small model and moves it up. On-policy distillation, used in Kimi K3 and DeepSeek-V4, still samples from the student and updates it. If the student is the expensive model, the rollout bill stays. Allspark asks a narrower question. Can a weak teacher trained with RL help a strong student when training never sees that student's rollouts, and test-time contact is only shared text?

Method

Training uses two copies of the same small model. The teacher π is updated with RL. The student ρ is a frozen clone of the initial weights; it continues the trace and writes the final answer. They take turns appending chunks to one chain of thought. The student speaks first. When reasoning stops, the student answers. Loss lands only on teacher tokens. Student tokens and boundary markers are masked. The reward is binary correctness of the final answer. There is no process reward on the trace. Freezing the training partner is how the teacher learns to help a model that will not update, which is the same constraint it will face with a frozen strong student.

A teacher chunk is useful if the student's continuation after it is more likely to be correct. That quantity is the continuation value Q(h, z): expected reward given prefix h and teacher span z. The teacher does not have to solve the problem. It has to say something the partner can use, like the weaker voice in pair programming.

At transfer the teacher is frozen and the training partner is swapped for a stronger student. Both stay frozen. Each side decodes and re-encodes the shared text with its own tokenizer, so architecture and vocabulary need not match.

The Qwen study trains LoRA on Qwen3-1.7B. The larger study trains Inkling-Small (276B total, 12B active per token) and transfers to Inkling (975B total, 41B active), then to Nemotron-3-Ultra and Kimi-K2.6.

Results

Qwen development accuracy, 128 questions per domain, one sample each, Allspark teacher at checkpoint 94:

StudentTeacherMathReasoning
Qwen3-1.7BNone81.373.4
Qwen3-1.7BOrdinary RL88.271.8
Qwen3-1.7BAllspark89.875.0
Qwen3-1.7B (RL-tuned)None89.870.3
Qwen3-4BNone96.981.2
Qwen3-4BUntrained93.079.7
Qwen3-4BOrdinary RL89.879.6
Qwen3-4BAllspark96.182.8

Same-size Qwen3-1.7B, Allspark beats both the solo student and an ordinary-RL teacher on both domains. Move to Qwen3-4B and naive teachers hurt: an untrained 1.7B teacher drops math from 96.9 to 93.0 and reasoning from 81.2 to 79.7; an ordinary-RL teacher drops them to 89.8 and 79.6. Allspark is roughly flat on math (96.1 vs 96.9) and up 1.6 points on reasoning. Parking a weak model in front of a strong student is a tax by default. The teacher has to be trained for the handoff. In selected 4B traces the teacher recovers an octagon area from 1053/2 to 567, and in another case it feeds a broken tens table that the student copies.

The larger study uses ARC-AGI-2. One thousand public training tasks split 904/96; the 96-task development panel uses three attempts each. Inkling alone peaks at 78.1%. Allspark peaks at 81.6%. At one operating point Allspark hits 79.2% with 10.8k retained tokens, above the solo student's best 78.1% at 12.3k.

The same frozen Inkling-Small teacher is then paired with other families. The paper reports deltas only, not absolute accuracies:

StudentAccuracy changeRetained tokens
Nemotron-3-Ultra medium+17.7 pp+72.5%
Nemotron-3-Ultra full+5.2 pp+4.4%
Kimi-K2.6 (128K)+11.5 pp-7.7%

An earlier within-family control (checkpoint 35, effort 0.7, one attempt) points the same way: an untrained teacher drops Inkling from 75.0% to 66.7%; the Allspark teacher reaches 76.0%.

Cost is worse than the token plot. At matched Inkling thinking effort, modeled cost per attempt is 1.42-1.82x the student alone. The highlighted Inkling comparison uses 12.1% fewer retained tokens (10,771 vs 12,253) but only about 2.7% less modeled cost ($0.0606 vs $0.0623). On Kimi, tokens fall 7.7% while cost rises 21.1% (about $0.1466 to $0.1774). Dual prefixes, cached rereads, and the teacher's sampling rate eat the token win.

Why it matters

For a lab that cannot run RL on the target model, this is a reusable inference plugin. Train the teacher once on small-model rollouts, freeze it, and pair it with a strong student you can already serve. The student's weights never move. Cross-tokenizer transfer is text.

It is not a free upgrade. Qwen3-4B math does not beat the student alone. ARC gains shift with thinking effort. Token savings do not become dollar savings. The student does not absorb the skill; the teacher has to stay in the serving path.

Limitations

The paper lists three. Inference serves both models, so extra calls and handoffs add latency and cost; fewer retained tokens do not imply cheaper deployment. The model, task, and run coverage is narrow. The implementation needs interrupt-and-resume on a shared reasoning stream; non-thinking models and APIs that hide that stream are untested.

A few more cracks. Qwen numbers come from the development panel used to pick checkpoints, one seed per method, not a public test set. ARC is 96 development tasks carved from the public training split, not the official test. In the cross-family comparison, Allspark explicitly closes the reasoning channel and asks for a separate answer; the solo student must switch inside one completion, so the protocols are not matched. In a coding case the teacher injects a broken tens table and the student copies it. Nemotron medium's +17.7 points arrives with +72.5% tokens; the paper does not separate extra length from extra usefulness.

Terms

Source

Related papers

All paper explainers