RULER: instance-aware rubric rewards beat reward hacking in SVG generation RL

inclusionAI · hf · 2026-09-23

Researchers introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning) for open-ended SVG generation, where scalar metrics like CLIP/Aesthetic scores transfer poorly and trigger reward hacking when used as RL rewards.

Approach: each instruction is converted into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and weighted satisfactions form a fine-grained reward optimized with GRPO. Since the rubric comes from text alone, no paired SVG ground truth or human preference labels are needed.

Results: rubric scores rise from 0.432/0.395 to 0.693/0.683 on MMSVG-Illustration/Icon, surpassing dedicated SVG specialists and matching the much larger DeepSeek-V3; ablations identify rubric design as the active lever.

Original post →

More from Multimodal

Multimodal channel →