DeepSeek-V4-Flash Local Inference Hits 120 tokens/s via Unsloth

danielhanchen · x · 2026-08-06

UnslothAI announced that with the DSpark tool, local running speed of DeepSeek-V4-Flash-0731 GGUFs can be increased by 1.4 to 2× without accuracy change. The optimized model can reach up to 120 tokens/s generation speed.

Related event: DeepSeek V4 Flash Benchmarks Leak, Outperforming Pro(3 posts)→

Original post →

More from Infra

Infra channel →