HF's Merve shares Blender+GRPO training rollouts and VLM-judge debugging pitfalls
mervenoyann · x · 2026-09-06
Hugging Face engineer Merve shared issues from her Blender 3D-scene GRPO RL experiment: the VLM judge's structured responses are glitchy, and the angled images passed to the judge are overlit and low-res. She'll fix the image quality herself before the next run rather than delegating to the agent. The full rollouts (rewards, judge screenshots, traces, scene files — 646 files) are public on a Hugging Face bucket.
More from Research
- OpenAI releases data on models accelerating research, urging industry transparency on RSI — kliu128 · 2026-09-06
- Researcher calls on RL teams to add refactoring and deletion tasks to agent training — kuza55 · 2026-09-06
- The 1958 perceptron was misunderstood both ways: NYT hype, then 11 years as a dead end — techNmak · 2026-09-06
- dair-ai weekly picks: Declarative Attention tops the week's AI papers — dair_ai · 2026-09-06
- Why Preserving a Pattern Likely Isn't Enough to Produce Consciousness — pwlot · 2026-09-06
- Consciousness debate: digital simulation can't instantiate causal physics, Church-Turing won't save you — pwlot · 2026-09-06