MEND: RL for flow models via proximal velocity matching beats Flow-GRPO in 100 vs ~4k updates
UTEXAS · hf · 2026-10-07
UT Austin researchers introduce MEND, a reinforcement learning method for reward post-training of flow models built on proximal velocity matching.
Method
- Caps rewards within each prompt group so well-scoring samples receive no update.
- Below the cap, proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price.
- Regresses onto resulting velocity targets with no KL term, frozen reference model, or advantage weights.
Results
- In 100 updates, MEND outperforms Flow-GRPO (4k updates) on five of six evaluators at the same distance to base-model images.
- Under equal budget, it surpasses ReFL and DiffusionNFT at every evaluated update across four rewards, reaching PickScore 24.03 vs 23.92 and 23.43.
- A 300-update three-reward run beats the five-reward DiffusionNFT model on all three rewards it trains on.
The method is general and applies to any flow backbone with a differentiable reward.
More from Multimodal
- Blender's The City Generator brings Houdini-style procedural 3D city building to everyone — bilawalsidhu · 2026-10-07
- MiniCPM-V-4.7-35B-A3B quietly appears on Hugging Face without a model card — BreakfastFriendly728 · 2026-10-07
- E-commerce product showcase animation built with Opus 5.5 + Higgsfield — Tegadesigns · 2026-10-07
- Gemini Omni 1.1 Can Be Surprisingly Cohesive When Used Right, Demo Shows — Cold_Solder_ · 2026-10-07
- EmbeddingGemma 2 Demo: One Search Across Images, Video and Audio in Browser — tomaarsen · 2026-10-07
- Creator builds 4-minute AI sci-fi trailer MORO with Agent Two and AI-assisted sound design — LudovicCreator · 2026-10-07