DiffusionOPSD: On-Policy Self-Distillation in Diffusion Models
ByteDance-Seed · hf · 2026-08-26
ByteDance proposed DiffusionOPSD, which uses on-policy self-distillation to turn image-level rewards into explicit intermediate targets for diffusion models. This approach improves alignment efficiency and enables separate analysis of target construction and policy fitting.
More from Research
- Catching bugs in scikit-learn by comparing versions — Lost-Dragonfruit-663 · 2026-08-26
- Gemini Flash 3.7 Passes Enterprise Agent Safety Benchmarks with GraphJin — dosco · 2026-08-26
- Roboticist Reflection: Prioritize Inference Behavior Over Model Training — deepakpathak · 2026-08-26
- Honesty about fake environments prevents model hallucinations — Sauers_ · 2026-08-26
- Study finds LLMs susceptible to 'Prior-hacking', derailing reasoning — RexDouglass · 2026-08-26
- SemaPLC: verification-gated agent loop nearly doubles dynamic behavior scores for AI-written PLC code — 量子位 · 2026-08-26