Flux-OPD: Solving LLM Distillation in Open-Ended Domains with Evolving Contexts

ZenMoore1 · x · 2026-08-01

Training LLMs in open-ended domains lacks verifiable rewards, making preference supervision difficult. This paper introduces Flux-OPD, an On-Policy Distillation (OPD) paradigm using evolving contexts.

The method dynamically updates the teacher's guidance context based on the student's training state. By decomposing the reverse KL objective, it uses a conflict term to weight contextual corrections injected into a context-free teacher anchor, resolving instability and distribution conflicts.

Experiments show Flux-OPD outperforms existing OPD paradigms in tasks like video prompt optimization and medical QA.

Original post →

More from Research

Research channel →