OLIVE: Teacher Continuations at Student Prefixes Beat Offline SFT by 13% on ScienceWorld

Learning from Teacher Continuations at Student States

Haojin Wang, Dylan Zhang, Huaibo Chen, Suhao Yu, Yihang Sun, Zhanyang Jin, Jiaying Ye, Dianqi Li, Prasanna Sattigeri, Kamal Youcef-Toumi, Hao Peng

cs.CL

2026-09-29

OLIVE trains on teacher continuations of student prefixes: +6–8 pass@8 on hard RLVE vs the base model, and +13% vs offline SFT on ScienceWorld with GPT-5.4-mini text.

What problem this solves

Two common ways to distill a stronger language model into a weaker one: train on frozen teacher trajectories (offline SFT), or run on-policy distillation (OPD), where the student rolls out and the teacher scores every token with reverse KL.

Offline SFT trains under teacher prefixes and then infers under the student's own tokens. That sequential covariate shift lets errors compound on long traces, and a fixed dataset cannot track the states a moving student actually visits. OPD does condition on student states, but later targets still sit on the student's original continuation. If the teacher would have flipped an earlier token, nothing in the rollout shows what should follow that flip. The picture gets worse when the capability gap is large, or when an early agent action leaves the teacher's support. Distribution matching also needs teacher logits, so teachers that expose only text cannot be used.

OLIVE puts the supervision on states the student actually reaches, and has the teacher show how to continue from there.

Method

OLIVE (OnLine InterVEntion) does three things each step:

Prefixes are redrawn after every update, so supervision tracks the moving policy, in the spirit of DAgger. Because the teacher continues under its own choices, the suffix can demonstrate a recovery path. The loss uses text only, so tokenizers may differ and the teacher can be a black box. Continuations may be truncated and unverified, which caps the cost of hosting the teacher online.

Default reasoning setup: k=4096, M=1024, with the per-rollout distilled-token budget matched to baselines at 7168. For longer agent environments (ALFWorld, ScienceWorld) the student plays 10 turns and the teacher at most 5; shorter environments use a 5-turn student prefix. Prefixes are kept unfiltered.

The synchronous version idles the student GPU until the teacher finishes. The async version overlaps the next batch of prefixes with the current teacher continuations, and caps staleness with an asynchronous depth d, the maximum number of student updates between sampling a prefix and using its trace. Reasoning runs use d=3.

Results

Teacher: Qwen3-4B-Thinking-2507. Testbed: 18 hard synthetic reasoning environments from RLVE, 9K training problems, thinking mode on. The OPD baseline uses a top-16 KL approximation.

StudentMethodPass@8Avg@8
Qwen3-1.7BBase11.13.3
Qwen3-1.7BOffline teacher traces15.05.6
Qwen3-1.7BOPD14.44.6
Qwen3-1.7BOLIVE19.47.8
Qwen3-4BBase46.118.1
Qwen3-4BOffline teacher traces53.319.9
Qwen3-4BOPD45.621.3
Qwen3-4BOLIVE52.223.5

OLIVE posts the largest avg@8 gains: +4.5 on 1.7B, +5.4 on 4B. On 4B pass@8, unfiltered offline teacher traces still edge it, 53.3 versus 52.2. During OPD, student-teacher top-K overlap barely moves, about 0.707 to 0.713, which matches the claim that token-level scoring stalls when the two models think differently.

On the same 9K set with 8×H200: OPD 41.3 GPU-hours, sync OLIVE 39.1, async OLIVE 29.8. Async cuts wall time 23.8% versus sync, with pass@8 falling from 54.4 to 52.2. Versus OPD it uses about 28% fewer GPU-hours and is 6.6 pass@8 points higher.

Agentic setting: Qwen3-1.7B student, Qwen3-32B teacher, avg@4 success.

MethodALFWorldScienceWorldTextCraftBabyAISearchQA
Student19.380.1223.0038.3330.50
OPD22.250.0029.5043.0629.56
Guided OPD27.120.7545.5067.7837.50
OLIVE40.007.5055.2567.5039.06

ScienceWorld starts near zero and OPD stays there; OLIVE reaches 7.5%. Guided OPD slightly leads on BabyAI, 67.78 versus 67.50. OLIVE leads on the other four environments.

With text-only GPT-5.4-mini, OPD is unavailable. Offline SFT on ScienceWorld plateaus after two epochs. OLIVE is slower across those two epochs and then finishes 13% higher at epoch 5. Average drop on general benchmarks (AIME25, LiveCodeBench v6, IF-Eval, GPQA Diamond): offline SFT −3.4, offline prefix continuation OEC −1.3, OLIVE −0.9. After sequential training on five agent environments, the last one, ScienceWorld, shows a 13.0-point gap in OLIVE's favor.

Why it matters

The SFT objective moves from frozen teacher traces to teacher continuations at states the current student visits. The loss is still cross-entropy and the interface is still text, so an API model can serve as an online coach without logits, and without the student first producing a full trajectory the teacher can score.

That is most useful when the student cannot yet write a reliable long chain of thought, or when an early agent mistake knocks the episode off the teacher's distribution. Against OPD, OLIVE shows the next steps. Against offline SFT, refreshed prefixes delay the plateau and forget less. Later tasks remain learnable.

The cost is hosting the teacher during training. Async brings that cost in line with OPD, and in this run it is lower.

This is an incremental move on where supervision is placed and whether it is refreshed. The paper does not compare against reward-based RL.

Limitations

The manuscript is marked Ongoing work and has no standalone Limitations section.

Async trades stale prefixes for throughput; at d=3 the 4B pass@8 drops 2.2 points. k and M are fixed, prefixes are unfiltered, and a broken prefix may be unrecoverable even for the teacher. None of these choices is swept. ScienceWorld at 7.5% is a jump from the student and still about half of the teacher's 15.62%, far from a usable agent. 4B reasoning pass@8 still trails unfiltered offline teacher traces. The main teacher is same-family Qwen3; the only cross-family black-box result is GPT-5.4-mini on ScienceWorld. There is no head-to-head with RL post-training. RLVE is contamination-free synthetic data; transfer to natural exam benchmarks is not reported.

Terms

Source

Related papers

All paper explainers