Image models can't count circles: Verifiable Visual Rewards boost instruction following

StellaLisy · x · 2026-09-29

Stella Lisy's team shows image generation models still fail at precise instruction following (object counts, spatial relations) and introduces Verifiable Visual Rewards (VVR): prompts and deterministic verifiers derived programmatically from geometric scenes, replacing unreliable reward models.

Key points:

Related event: Verifiable Visual Rewards Boost SD3.5 Instruction Following from 2.8% to 28.3%(3 posts)→

Original post →

More from Multimodal

Multimodal channel →