HF's Merve shares Blender+GRPO training rollouts and VLM-judge debugging pitfalls

mervenoyann · x · 2026-09-06

Hugging Face engineer Merve shared issues from her Blender 3D-scene GRPO RL experiment: the VLM judge's structured responses are glitchy, and the angled images passed to the judge are overlit and low-res. She'll fix the image quality herself before the next run rather than delegating to the agent. The full rollouts (rewards, judge screenshots, traces, scene files — 646 files) are public on a Hugging Face bucket.

Original post →

More from Research

Research channel →