DeepSeek V4 Flash fits on a 128GB Ryzen AI MAX+ 395 and hits 32 tok/s
sandropuppo · reddit · 2026-07-28
A Reddit post shows how DeepSeek V4 Flash was fit onto a single AMD Ryzen AI MAX+ 395 system with 128 GB of unified memory, reaching up to 32 tok/s with speculative decoding.
Key details:
- The model was compressed with the ROCmFPX family of block formats, using a mixed-precision recipe that brought the target to about 102.3 GB, or roughly 2.88 bits per parameter.
- With no speculative draft, the DeepSeek-specific HIP decode path reached 25.31 tok/s autoregressively.
- Adding a small DSpark draft and a q=4 verification cap pushed throughput to 32.0 tok/s.
- Sparse prefill reached roughly 245 tok/s in the published run, with separate validation cases ranging from 246.8 to 255.9 tok/s around the 8K context window.
The post also notes that sparse prefill is opt-in because its reduction order is not byte-identical to exact tokenwise prefill, even though small smoke tests passed.
More from Infra
- YC startup hwintelligence launches Wave, a waveform debugger with a verification agent — ycombinator · 2026-07-29
- Broadcom says AI could cut exploit windows from weeks to hours — therealdanvega · 2026-07-29
- camelAI moved its agent from VMs to a Cloudflare Durable Object — irvinebroque · 2026-07-29
- LLM Inference Handbook collects deployment, GPU, and optimization guidance for production teams — carrycooldude · 2026-07-29
- MLCommons launches MLPerf Endpoints v0.7 for AI inference benchmarking — TheKanter · 2026-07-29
- Solo project compares cloud providers and generates Terraform — Character-Ring5785 · 2026-07-29