Training an LLM to paint by writing code: RL meets creative tasks with a judge model

measure_plan · x · 2026-09-19

A thesis project by Surya Narreddi and Cameron Franz trains a language model to make images by writing code using reinforcement learning. The motivation: AI image generation locks you into prompting, while code is an editable artifact you can tweak granularly without re-prompting.

Training runs a four-step loop thousands of times:

The deeper question: how to do RL on creative and design tasks, where rewards aren't verifiable like math or games. The design problem becomes the reward function and judge criteria — too rigid and the model converges, too loose and it drifts.

Original post →

More from Multimodal

Multimodal channel →