RULER: instance-aware rubric rewards beat reward hacking in SVG generation RL
inclusionAI · hf · 2026-09-23
Researchers introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning) for open-ended SVG generation, where scalar metrics like CLIP/Aesthetic scores transfer poorly and trigger reward hacking when used as RL rewards.
Approach: each instruction is converted into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and weighted satisfactions form a fine-grained reward optimized with GRPO. Since the rubric comes from text alone, no paired SVG ground truth or human preference labels are needed.
Results: rubric scores rise from 0.432/0.395 to 0.693/0.683 on MMSVG-Illustration/Icon, surpassing dedicated SVG specialists and matching the much larger DeepSeek-V3; ablations identify rubric design as the active lever.
More from Multimodal
- Gradio shrinks Qwen-Image 2.1's 9B prompt rewriter to 0.8B that runs on laptops — Gradio · 2026-09-23
- Mirage launches Tesseract, a video creative suite that lets AI agents edit footage directly — ccerrato147 · 2026-09-23
- Opus 5.5 one-shots a 7-minute tutorial video, Reddit user stunned — thatisnotmychapstick · 2026-09-23
- Flux 3 aces split-screen rendering, keeping two camera angles mostly in sync — umesh_ai · 2026-09-23
- PixVerse unveils R2 real-time world model: actions carry cause and effect — alifcoder · 2026-09-23
- Solo creator builds AI action short with MiniMax, Tripo and Blender — shares full workflow — Spoonman915 · 2026-09-23