Teknium: DeepSeek V4 Flash Hits 70 tok/s on Local Inference

Teknium · x · 2026-08-03

Prominent open-source model author Teknium stated during a discussion on local LLM inference hardware that running the DeepSeek V4 Flash model on devices like DGX Spark can achieve up to 70 tokens per second.\n\nThis directly addresses community concerns about the inference performance of expensive local hardware, proving that specific lightweight models already possess highly practical throughput rates on edge devices.

Original post →

More from Infra

Infra channel →