Distilling DeepSeek V4 Flash to a 4B model on DGX Spark: 26 hours, 22ms per judgment

Dan_Jeffries1 · x · 2026-09-19

Developer @taroleo shares a local distillation experiment: spending 26 hours on a DGX Spark to distill the judgment capability of the 157GB-weight DeepSeek V4 Flash into a 4B model.

Key takeaways:

The insight: instead of having a big model emit full JSON, strip generation down to classification-only and gain order-of-magnitude speed. A useful reference for distillation + edge deployment.

Original post →

More from Infra

Infra channel →