DeepSeek-V4-Flash on 4×AMD V620: 300K Context, 21 tok/s Generation
Thin_Pollution8843 · reddit · 2026-08-15
A user runs DeepSeek-V4-Flash-0731 IQ3XXS on 4 AMD Radeon Pro V620 GPUs (128GB VRAM) with DSpark speculative decoding. 32K prompt ingestion at 276 tok/s, short prompt at 379 tok/s, continuous generation at 21 tok/s, accepted generation at 30.6 tok/s. However, q3xxs quantization degrades quality, and VRAM is insufficient for the drafter.
More from Infra
- NVIDIA open-sources NeMo Switchyard for dynamic model routing in agent workflows — NVIDIAAI · 2026-08-15
- Vercel ranked as the world's fastest AI Gateway infrastructure — cramforce · 2026-08-15
- Qwen3.8-2.4T-A95B deployment guide: NVFP4 needs 8×B300, TP must divide 64 — Necessary_Gazelle211 · 2026-08-15
- mcpp: Auto-generate MCP servers from C++ code via reflection — karurochari · 2026-08-15
- RTX 3090 gets 35 t/s on Qwen 3.8 27B — cviperr33 · 2026-08-15
- CME to launch futures contracts tracking Nvidia H100/B100 compute costs — AccBalanced · 2026-08-15