Full-precision DeepSeek 4.1 Flash hits 300+ TPS on 4 RTX Pros with custom vLLM fork
TheZachMueller · x · 2026-09-11
A developer (@fraserpricee) demonstrated running full-precision DeepSeek 4.1 Flash at 300+ TPS on four RTX Pro cards, with under 32GB peak system RAM, using a custom vLLM fork and a small SSD. The full recipe is promised to be published later; for now it's a teaser of the local inference setup.
More from Infra
- AGI as task time horizon vs meetings — and why fabs should train their own models — jwt0625 · 2026-09-11
- Free client-side calculator compares LLM token economics across DeepSeek, Claude, o3-mini — nikola_mr64990 · 2026-09-11
- 10 resources on what happens after training: KV-cache, quantization, serving — techNmak · 2026-09-11
- Inside a Trillion-Token/Day Factory: Mooncake Turns KVCache into a Cluster-Shared Pool — 量子位 · 2026-09-11
- Web Demo Approximates V4.1 Flash-Style Fast KV Prefill on Qwen3 — T_rex2700 · 2026-09-11
- OpenAI CFO: Compute I bought a year ago could sell for 3-5x today — and we're still short — rwang07 · 2026-09-11