IR4RL turns intermediate render progress into RL rewards, new SOTA for image-to-code

phillip_isola · x · 2026-10-07

Researchers including Phillip Isola and Tali Dekel introduce IR4RL (Reinforcement Learning from Intermediate Renders) for RL post-training of image-to-code VLMs. Standard outcome-only RL computes rewards from the final render, giving sparse feedback misaligned with individual tokens — a program can be partly correct and partly wrong yet all tokens share one terminal reward. IR4RL instead converts changes between intermediate renders into token-level render-progress rewards, yielding denser localized supervision. On Image-to-SVG and Image-to-TikZ it beats supervised fine-tuning and standard GRPO, achieving new state-of-the-art. Project page with interactive demo; models on Hugging Face, code coming soon.

Original post →

More from coding & agent

coding & agent channel →