WTF paper fine-tunes flow maps with optimal transport, cutting RL training compute up to 280x
yeewhye · x · 2026-09-27
A new paper by Abbas Mammadov, Yee Whye Teh, Nicholas Boffi et al. introduces Wasserstein-Tilted Flow Maps (WTF), a simulation-free RL recipe for reward fine-tuning of flow-based generative models.
- Existing methods treat reward fine-tuning as sampling from a reward-tilted distribution via KL regularization, which merely reweights the base distribution
- WTF instead uses an optimal transport regularizer built from the pre-trained drift, transporting individual samples toward higher-reward regions, and shows this is equivalent to a deterministic optimal control problem on the flow
- This yields the first end-to-end fine-tuning recipe native to flow maps: the output retains strong reward-aligned performance at few-step inference budgets without post-hoc distillation
- On ImageNet-256 and text-to-image experiments, WTF achieves higher reward with comparable or higher diversity than baselines while using up to 280x less training compute
More from Research
- A 'new type of mind' trained without natural language writes kernels at up to 66x Triton speed — repligate · 2026-09-27
- ENGRAM and YOCO memory architectures could shift data center DRAM/HBM economics — AccBalanced · 2026-09-27
- From Pixels to Visual Tokens: How LLMs Actually 'See' Images — makaros622 · 2026-09-27
- Spatial-Interactor teaches VLMs spatial reasoning through physical interaction — arankomatsuzaki · 2026-09-27
- DeepMind researcher explains why AI hasn't transformed physics yet — DaniloJRezende · 2026-09-27
- IIT Delhi lands 4 NeurIPS 2026 main-track papers on multilingual interpretability and distillation — Tanmoy_Chak · 2026-09-27