JD's Any-OPD Enables Cross-Architecture Distillation for Flow-Matching Image Models
jingdong1 · hf · 2026-08-05
A JD research team proposed Any-OPD, a framework solving the on-policy distillation problem for flow-matching models with mismatched architectures.
Background & Challenges
- Standard on-policy distillation requires identical VAE latents, matching architectures, and a common timestep grid between teacher and student.
- When the strongest teacher and the deployable student come from different families, standard recipes fail: teacher latents cannot serve as targets, and per-pixel losses lead to blur or divergence.
Any-OPD Solution
- Treats the teacher purely as a black-box sampler, connecting the two models via a frozen, model-agnostic vision representation space.
- Recovers trajectory correspondence by matching continuous noise levels instead of step indices.
- Introduces a brief anchoring phase to ensure the on-policy gradient measures sample quality rather than domain mismatch.
Results
- Distilled the 12B FLUX.1-dev into the 2.5B SD3.5-Medium.
- Lifted the student's PickScore from 0.846 to 0.884 and HPSv3 from 9.12 to 10.97.
- Rivaled the teacher's quality at a fifth of the size.
More from Multimodal
- What Makes an AI Music Video Feel Like a Real Video? — ThemeOld5001 · 2026-08-05
- AI Agent Autonomously Produces Mini Documentary End-to-End — illscience · 2026-08-05
- Open Source Community Slashes MiniMax H3 Video Model VRAM to 5GB in 48 Hours — ostrisai · 2026-08-05
- Alibaba's Qwen3.8-Max Takes #2 Spot on Image-to-WebDev Arena — rohanpaul_ai · 2026-08-05
- Minimax H3 Generates Crossover Video: Seinfeld Meets FRIENDS — Time-Ad-7720 · 2026-08-05
- AI Video Imagines Crossover Date Between Friends and Seinfeld — Time-Ad-7720 · 2026-08-05