Distilled 260M text-to-image model hits <190ms on a T4, generating per keystroke

Gradio · x · 2026-10-01

Gradio details the ML-Intern distilled 260M text-to-image model: distillation cut network passes per image from 100 to just 4, bringing generation under 190 milliseconds on a T4 GPU — fast enough to render an image on every keystroke. It runs on a small GPU Space or locally in the browser via WebGPU.

Original post →

More from Infra

Infra channel →