Distillation Costed ~$16 Total: It's a Data-Cost Problem, Not a Training-Cost Problem
Gradio · x · 2026-09-23
Gradio broke down the bill for its rewriter distillation: $16 total. Training both models: 30 minutes and $0.75 on one A10G. Making labels: 2h37m on an A100, $6.54. Evaluating with 160 rendered images: $3.45. 80% of the cost went to data and evaluation — model distillation is a data-cost problem, not a training-cost problem. Two Gradio Spaces ship with it: Pocket Studio (any-language input, 0.8B/2B, editable rewrite, render) and Rewriter Arena (same-seed side-by-side of 9B/0.8B/2B/raw plus 40 pre-rendered comparisons).
Related event: Gradio's Distillation Project Cost Just $16, Run End-to-End by AI(2 posts)→
More from Infra
- Dev laments agents built around KV caches, wants inference-first chips — dbreunig · 2026-09-23
- GE Vernova seen hitting $200B backlog by early 2027 as turbine demand outruns guidance — BenBajarin · 2026-09-23
- New LLM papers: recursive language models generalize out of domain; XMerge depth compression — burny_tech · 2026-09-23
- DeepSeek report praised: absurdly tiny KV cache, dense infra design — stochasticchasm · 2026-09-23
- Running Qwen 27B locally on RTX 4090: beats pre-2025 coding models, RAM is the wall — julianharris · 2026-09-23
- FP8 Tuning Cuts 42.9ms Per Step: Custom SGLang Kernels Boost Inference 126% — HankYeomans · 2026-09-23