Live human votes plugged into Flow-GRPO to stop image models gaming reward models
lmoroney · x · 2026-10-07
- Learned reward models like PickScore and ImageReward can drift from human judgment as a fine-tuned image model's outputs change, and training learns to exploit them.
- Rapidata's Hugging Face blog post proposes plugging live human votes directly into a Flow-GRPO training loop: the model generates a group of images per prompt, humans compare them pairwise, a Bradley-Terry fit converts votes into per-image scores, and those scores become group-relative advantages for the update.
- Rough math for the original Flow-GRPO setup (48 groups of 24 images per step, 100-150 responses per group, 3-minute deadline) implies 2,400 human responses per minute. Note the speed claims come from Rapidata's own product.
More from Multimodal
- Full Arabic Infographic Prompt: Define the Palette Before Rendering — aziz4ai · 2026-10-07
- Nano Banana 2.1 Infographic Test With Full Prompt Shared — aziz4ai · 2026-10-07
- Sarvam opens Content Studio dubbing APIs that keep each speaker's original voice — itsOmSarraf_ · 2026-10-07
- An 8-minute AI-made data video hit 3M views — here's the full workflow — pcuenq · 2026-10-07
- One prompt produced a hype-filled 'techno-optimist' music video — gerardsans · 2026-10-07
- AI short film 'The Unit' built with LTX2.5, Krea2 and Qwen Image 2.1 — No-Judge-2848 · 2026-10-07