Lightning Weave composes reasoning capabilities via on-policy distillation
Yecheng Wu · hf · 2026-09-16
Lightning Weave is a new method that composes independently trained reasoning capabilities into a single efficient student model via on-policy distillation, improving both accuracy and token efficiency on math and code benchmarks.
More from Research
- Composing Continual Learning Mechanisms Boosts Long-Horizon Memorization in LMs — JohnsHopkins · 2026-09-16
- Gavel Elicits Native Skill Routing from a Frozen LLM's Hidden States — Tsinghua · 2026-09-16
- ScienceBuddy: Recursive-in-Recursive Self-Improvement for Scientific Agents — Shuhan Xue · 2026-09-16
- RCT: LLMs fail to significantly boost novices' wet-lab molecular biology success — shae_mcl · 2026-09-16
- Six years on, scvi-tools still widely used — outlasting foundation models and coding agents — anshulkundaje · 2026-09-16
- Danish administrative data finds no effect of AI on earnings or wages — paulnovosad · 2026-09-16