FULL STORY
Redis Author Optimizes DeepSeek V4
Redis author antirez shared progress on the DwarfStar engine, significantly boosting DeepSeek V4 inference speeds to 170 t/s via quantization and routing.
2026-08-17 ~ 2026-08-17 · 2 episodes · 5 posts
Episode 1 · Redis Author Optimizes DeepSeek V4 to 45 t/s (2026-08-17, 3 posts)
Redis author antirez demonstrated the DwarfStar inference engine, optimizing DeepSeek V4 PRO to 45 t/s via Q2 quantization and dynamic expert routing. The system adaptively migrates experts between VRAM and RAM based on usage.
- DeepSeek v4 PRO Q2 hits 45 t/s with VRAM/RAM split on DGX Station — antirez · 2026-08-17
- antirez's DwarfStar: adaptive VRAM/RAM expert placement runs DeepSeek V4 at 45 t/s — antirez · 2026-08-17
- Dynamic Expert Migration: Adaptive Optimization from VRAM to RAM — antirez · 2026-08-17
Episode 2 · antirez Tunes DwarfStar to 170 t/s Generation on Station (2026-08-17, 2 posts)
Redis creator antirez reported that after roughly 48 hours of optimization, the DwarfStar inference engine on Station achieved 170 t/s generation (without DFlash) and 22k tokens/s prefill using MXFP4 and Flash models.
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17
- DwarfStar engine hits 170 t/s on Station with MXFP4 optimization — antirez · 2026-08-17