Verifiable Visual Rewards lift SD3.5 instruction accuracy from 2.8% to 28.3% on arXiv
testingcatalog · x · 2026-09-29
Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, and Luke Zettlemoyer released the arXiv paper "Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts."
- Problem: precise instruction following (object counts, spatial relations) in image generation remains open, partly because training relies on unreliable reward models (object detectors, VLMs)
- Method: VVR is the first programmatically verifiable image reward framework — each task is a geometric scene from which prompt and deterministic verifier are auto-generated at any scale and complexity
- Data: releases VVRBench (10,000 tasks, 32 constraint types) and VVRBench-Challenge (720 harder tasks); strongest evaluated model GPT-Image-2.5 solves just 21.4%
- Results: RLVVR raises Stable Diffusion 3.5 Medium's VVRBench accuracy from 2.8% to 28.3% with consistent easy-to-hard and out-of-domain generalization; mixing VVR into existing objectives further improves quality and human preference, motivating adoption in standard post-training recipes
33-page paper with code and benchmark open-sourced.
More from Multimodal
- NVIDIA's LongLive-Plug: Distill Once, Deploy Training-Free Across 54 Downstream Video Models — nvidia · 2026-09-30
- Adobe Research Shows Adversarial Post-Training Restores Missing High-Frequency Detail in Pixel Diffusion — adobe-research · 2026-09-30
- UCSD's LIFT Lets You Control Future Video Layouts via On-Policy Self-Distillation — UCSanDiego · 2026-09-30
- MiniMax-H3 RefMod Upgrade: Packing JPEGs Into Safetensors Cuts Encoding Time and Kills Identity Bleed — acedelgado · 2026-09-30
- Creator Says Opus 5.5 Renders in 45 Mins What Used to Need Four-Figure Plugins — justin_hart · 2026-09-30
- Casey: AI video will inevitably get so good that critics must admit great works — johncoogan · 2026-09-30