E2S finetuning turns off-policy expert data into on-policy student samples via amortized sampling
Lianhuiq · x · 2026-10-06
Researcher haoqik322 shares E2S (Expert-to-Student) Finetuning, a weekend project presenting an amortized sampling approach that converts off-policy expert demonstrations into more on-policy SFT data — "make SFT great again."
- Problem: vanilla SFT forces the student to imitate correct but off-policy expert trajectories, causing unnecessary policy shifts and forgetting.
- Idea: train the model to "speak expert data in the student's own language," formulated as constrained generation — the expert defines what privileged information to preserve, the student decides how to generate under that constraint.
- Relation to prior work: Finetuning with Sampling targets this distribution with per-example MCMC, but inference-time search per example is expensive; E2S amortizes the search by learning a conditional sampler across examples.
Advisor Lianhui Qin amplified the thread.
More from Research
- AI math proofs shift to open-source-style collaboration, square packing shows — ctjlewis · 2026-10-07
- Overmind: Open Platform That Turns Production Traces Into Fine-Tuning Data for Agents — cneuralnetwork · 2026-10-07
- Study: Top-k Logits Leak as Much Information as Tuned Lens Trajectories, Far More Accessible — sineadwilliamso · 2026-10-07
- AutoAWQ Author: Reproduce Bonsai 2-Class Ternary Model for ~$43k on One B300 Node in ~4 Weeks — airesearch12 · 2026-10-07
- COLM 2026: Robust adaptation study unifies safety pretraining and midtraining — AdtRaghunathan · 2026-10-07
- SWE Decision Index v0.3 adds private benchmarks and vision evaluation — multimodalart · 2026-10-07