Verifiable Visual Rewards Boost SD3.5 Instruction Following from 2.8% to 28.3%

An arXiv paper by Shuyue Stella Li and colleagues proposes Verifiable Visual Rewards (VVR), lifting SD3.5's instruction-following accuracy from 2.8% to 28.3%. The verifier reliably distinguishes RL-generated outputs from baseline renders.

2026-09-29 ~ 2026-09-29 · 3 related posts