Latents are unevenly sensitive, so GRPO exploration should adapt per element
StellaLisy · x · 2026-10-08
GRPO in flow-matching models adds standard gaussian noise at each denoising step with an identical distribution per latent element. But latents are unevenly sensitive, so the model can end up exploring 'un-useful' directions — exploration should adapt to each element's sensitivity.
More from Research
- No human has read OpenAI's Navier-Stokes proof, says critic — formal verification alone isn't proof — gerardsans · 2026-10-09
- COLM paper: "forks in the road" in post-training data shrink reasoning model coverage — ParshinShojaee · 2026-10-09
- TermGrade: 1k open-source executable terminal-agent RL environments with full training recipe — maximelabonne · 2026-10-09
- tangermeme, a toolkit for interpreting cis-regulatory logic, published in Nature Methods — jmschreiber91 · 2026-10-09
- Google's open medical VLM MedGemma publishes in Nature Medicine — ymatias · 2026-10-09
- DuoMatching paper enables few-step video generation via joint-marginal distribution matching — Total-Resort-3120 · 2026-10-09