ByteDance Seed's RWTD lifts one-step SANA Sprint GenEval from 0.73 to 0.80
ByteDance-Seed · hf · 2026-10-02
ByteDance Seed introduces Reward-Weighted Transport Distillation (RWTD), a post-training method for one-step generative models requiring only generated samples and scalar reward evaluations. Rather than aligning to a conventional reward-tilted reference, RWTD builds an adaptive target mixing reward-tilted current and reference distributions — capturing in-training improvements while anchoring to the pretrained generator — realized via feature-space optimal transport and fixed-point regression. Theory shows the fixed-point interpolates between off-policy and on-policy reward tilting. Empirically it raises SANA Sprint 1.6B's GenEval from 0.73 to 0.80, with strong cross-reward generalization and preserved compositional capabilities.
More from Multimodal
- Magnific shares six Ideogram 4.5 posters with full prompt iteration threads — charis_ai · 2026-10-02
- Invoke 7 PR opens after 3 months and 320 PRs: new canvas engine, projects, video with audio — joshwcorbett · 2026-10-02
- Full Cinematic Prompt Breakdown: A Mountain-Scale Storm Dragon in Four Shots — ArjanDoge · 2026-10-02
- BFL opens free trials of FLUX Tools, commercial weights available on request — bfl_ai · 2026-10-02
- FLUX 3 Image just released, per early reports — chrisfirst · 2026-10-02
- MiniMax H3 recreates the 1991 cartoon Doug — and it turned out better than expected — Certain_Potato_4509 · 2026-10-02