SFT made it slower: OpenUI's self-training loop lifted generative-UI validity from 13% to 71.7%
1glasspaani · reddit · 2026-09-09
The OpenUI team details fine-tuning an open-weight model for Generative UI on consumer GPUs. Starting from DiffusionGemma, their first SFT attempt—a LoRA on 700 synthetic examples—lowered training loss but worsened benchmarks: broken component props and references. Constrained to a single component library, structural validity rose from 13% to 28.8%, but generation slowed from 1.6s to 4.3s with roughly double the denoising steps.
They then switched to a self-training loop:
- Generate programs with the current model
- Use the OpenUI Lang parser to accept valid outputs and pinpoint defects in the rest
- Repair only the defects, reject broad rewrites, and verify survivors against prompts with a judge
- Fine-tune on accepted examples and repeat
Result: 57.1% structural validity at 1.9s generation; the median repair changes one statement, and each pass takes 1–2 hours on a single A100. Repeating across 27 component libraries, the final model OUI-1 reaches 71.7% (structural validity only—not a guarantee of good looks or task completion). Blog, weights, and benchmark are public.
More from Research
- OpenAI claims agent swarm solved the Navier-Stokes Millennium Prize Problem — austinc3301 · 2026-09-09
- 88 hours to crack a 90-year-old fluid problem: OpenAI agents tackle Navier-Stokes — mdancho84 · 2026-09-09
- Report: OpenAI allegedly threatened NYU mathematician who declined co-authorship on Navier-Stokes solution — PsychicorAI · 2026-09-09
- Spilled.ink launches crowd-voted tracker for the OpenAI Navier–Stokes proof debate — NathanpmYoung · 2026-09-09
- DNA sequence models hit SOTA but wet-lab validation remains the bottleneck, experts say — anshulkundaje · 2026-09-09
- Researcher says he predicted Navier-Stokes would fall first, but not amid bitter controversy — geoffreyirving · 2026-09-09