DwarfStar engine hits 170 t/s on Station with MXFP4 optimization
antirez · x · 2026-08-17
Redis creator antirez shared performance benchmarks using the DwarfStar inference engine on Station. By optimizing for roughly 48 hours with MXFP4 and the Flash model, he achieved 170 tokens/sec generation and 22k tokens/sec prefill speeds. He also noted testing PRO and Q4 GLM 5.2 to evaluate how hardware-aware inference engines maintain performance when models exceed VRAM capacity.
Related event: antirez Tunes DwarfStar to 170 t/s Generation on Station(2 posts)→
More from Infra
- Bittensor Subnet 118 Adds Ultra-Cheap Inference, Joining Major AI Providers — markjeffrey · 2026-08-17
- Meta to rely on Nvidia Blackwell, AMD Helios in 2026, accelerate custom MTIA in 2027 — Beth_Kindig · 2026-08-17
- Stripe to Acquire OpenRouter for Over $7B, 5.4x May Valuation — rohanpaul_ai · 2026-08-17
- Wici One claims to solve local VRAM limits via NVMe offloading — Torodaddy · 2026-08-17
- Qwen3.8-27B hits 206 tok/s on single RTX 5090 via SGLang — StefanoGogioso · 2026-08-17
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17