Verifiable Visual Rewards Boost SD3.5 Instruction Following from 2.8% to 28.3%
An arXiv paper by Shuyue Stella Li and colleagues proposes Verifiable Visual Rewards (VVR), lifting SD3.5's instruction-following accuracy from 2.8% to 28.3%. The verifier reliably distinguishes RL-generated outputs from baseline renders.
2026-09-29 ~ 2026-09-29 · 3 related posts
- Image models can't count circles: Verifiable Visual Rewards boost instruction following — StellaLisy · 2026-09-29
- VVR paper demos: verifier cleanly separates RL outputs from baselines — StellaLisy · 2026-09-29
- Verifiable Visual Rewards lift SD3.5 instruction accuracy from 2.8% to 28.3% on arXiv — testingcatalog · 2026-09-29