ByteDance Seed's RWTD lifts one-step SANA Sprint GenEval from 0.73 to 0.80

ByteDance-Seed · hf · 2026-10-02

ByteDance Seed introduces Reward-Weighted Transport Distillation (RWTD), a post-training method for one-step generative models requiring only generated samples and scalar reward evaluations. Rather than aligning to a conventional reward-tilted reference, RWTD builds an adaptive target mixing reward-tilted current and reference distributions — capturing in-training improvements while anchoring to the pretrained generator — realized via feature-space optimal transport and fixed-point regression. Theory shows the fixed-point interpolates between off-policy and on-policy reward tilting. Empirically it raises SANA Sprint 1.6B's GenEval from 0.73 to 0.80, with strong cross-reward generalization and preserved compositional capabilities.

Original post →

More from Multimodal

Multimodal channel →