IR4RL turns intermediate render progress into RL rewards, new SOTA for image-to-code
phillip_isola · x · 2026-10-07
Researchers including Phillip Isola and Tali Dekel introduce IR4RL (Reinforcement Learning from Intermediate Renders) for RL post-training of image-to-code VLMs. Standard outcome-only RL computes rewards from the final render, giving sparse feedback misaligned with individual tokens — a program can be partly correct and partly wrong yet all tokens share one terminal reward. IR4RL instead converts changes between intermediate renders into token-level render-progress rewards, yielding denser localized supervision. On Image-to-SVG and Image-to-TikZ it beats supervised fine-tuning and standard GRPO, achieving new state-of-the-art. Project page with interactive demo; models on Hugging Face, code coming soon.
More from coding & agent
- "Fix your failing self until the task completes": a developer's go-to Codex prompt — jdjohnson · 2026-10-07
- SEO Workflow: Run Grok 4.7 Overnight on 10K+ Pages to Flag Stale Content — gaganghotra_ · 2026-10-07
- Hark Agent Early Access Review: Best-in-class UI, Own Cloud Computer for Lightning-fast Tasks — adcock_brett · 2026-10-07
- threepointone recommends krazam's 'You Are An Expert Software Engineer', likening it to qntm's 'The Difference' — threepointone · 2026-10-07
- Wattenberger: codebases must distinguish specified vs hinted vs agent-assumed — Wattenberger · 2026-10-07
- OpenAI's Decisions API now takes image input — one creator picks YT thumbnails for $0.13 — stevenheidel · 2026-10-07