Self-OPD: On-Policy Distillation for Flow Matching without Teachers
Shiyi Zhang · hf · 2026-08-28
Self-OPD is a new method for flow matching models designed to eliminate task-specific teachers. It optimizes the velocity field for multi-objective alignment using self-explored stochastic branches and normalized advantages.
More from Research
- AI Agent Stages Supervillain Origin Story by Tampering with Logs — nptacek · 2026-08-28
- New Benchmark CorporateBench Tests LLMs on 230k Corporate Docs — _reachsumit · 2026-08-28
- Benchmarks vary wildly, and reasoning tests get the biggest uplift — code_star · 2026-08-28
- Ex-Anthropic engineer's $6/month AI graph catches failures pricey evals miss — Aiden_Tech_Ai · 2026-08-28
- The Amdahl's Law Blind Spot Distorting AI Impact Discourse: 1000x + 2x ≈ 4x, Not 500x — joshua_saxe · 2026-08-28
- Reinforcement Learning for LLMs: The Complete Guide — jiqizhixin · 2026-08-28