One prompt, under $20, one night: how Gradio distilled a 4-step text-to-image model

Gradio · x · 2026-10-01

Gradio breaks down how ML-Intern built its text-to-image model: a single prompt on HuggingChat baked in distillation guidance, halving inference steps from 32 to 4 (32→16→8→4). The 4-step model was then fine-tuned with a perceptual loss and shipped only after passing a quality gate. Total cost: under $20 in a single overnight run, MIT-licensed.

Related event: Gradio Open-Sources 4-Step Text-to-Image Mini Model Trained for Under $20(2 posts)→

Original post →

More from Multimodal

Multimodal channel →