DeepSeek v4 PRO on DwarfStar: Peaks at 50 t/s with Dynamic VRAM/RAM
antirez · x · 2026-08-18
DeepSeek v4 PRO tested on DwarfStar running on a DGX Station, achieving 50 tokens/s peak without DSpark. Key optimizations include dynamic VRAM/RAM allocation for experts based on historical stats, 4k t/s prefill for large chunks, and layer-major streaming. The author notes that mixed RAM/VRAM inference requires careful handling.
More from Infra
- Running Qwen 27B at F16: Performance and VRAM Needs — Blues520 · 2026-08-18
- Qwen3.8-27B Benchmarks on M2 Ultra 192GB — planetearth80 · 2026-08-18
- GitNexus Boosts Coding Agent Performance by 30% — ycombinator · 2026-08-18
- Complete Guide to Enabling SageAttention on RDNA4: RX 9070 XT Tested — eloxH1Z1 · 2026-08-18
- Qwen3.8-27B optimization hits 1150 tps on RTX 3090 — iamMess · 2026-08-18
- RTX 4090 Config for Qwen 2.5 27B: No RAM Spill — gavwhittaker · 2026-08-18