MEND: RL for flow models via proximal velocity matching beats Flow-GRPO in 100 vs ~4k updates

UTEXAS · hf · 2026-10-07

UT Austin researchers introduce MEND, a reinforcement learning method for reward post-training of flow models built on proximal velocity matching.

Method

Results

The method is general and applies to any flow backbone with a differentiable reward.

Original post →

More from Multimodal

Multimodal channel →