Image models can't count circles: Verifiable Visual Rewards boost instruction following
StellaLisy · x · 2026-09-29
Stella Lisy's team shows image generation models still fail at precise instruction following (object counts, spatial relations) and introduces Verifiable Visual Rewards (VVR): prompts and deterministic verifiers derived programmatically from geometric scenes, replacing unreliable reward models.
Key points:
- Releases VVRBench (10,000 tasks, 32 constraint types) and harder VVRBench-Challenge (720 tasks); best model GPT-Image-2.5-Sunburst solves only 21% of the Challenge set
- Using VVR scores as RL rewards (RLVVR) lifts Stable Diffusion 3.5 Medium from 2.8% to 28.3% on VVRBench, with easy-to-hard generalization
- Mixing VVR into existing post-training objectives further improves overall quality and human preference
More from Multimodal
- VideoLoop rewrites bounded working memory, hits 88.3% on VideoMME long video — Jinfa Huang · 2026-09-30
- I used Codex to make a 15-second animation — and still had to give it editing notes — naridubs · 2026-09-30
- Flux 3 nails four-way split-screen images of one event from four angles — umesh_ai · 2026-09-30
- MageTrail 2.8B booru finetune costs $593 so far, hits limits of 41k-image dataset — Turbulent-Bass-649 · 2026-09-30
- Filmmakers jam with AI video generation wait times to shoot a music duet with Luma — mrjonfinger · 2026-09-30
- Meta's LSRM Wins ECCV 2026 Honorable Mention, Beats 3D Reconstruction SOTA by 2.4 dB — rsasaki0109 · 2026-09-30