DwarfStar engine hits 170 t/s on Station with MXFP4 optimization

antirez · x · 2026-08-17

Redis creator antirez shared performance benchmarks using the DwarfStar inference engine on Station. By optimizing for roughly 48 hours with MXFP4 and the Flash model, he achieved 170 tokens/sec generation and 22k tokens/sec prefill speeds. He also noted testing PRO and Q4 GLM 5.2 to evaluate how hardware-aware inference engines maintain performance when models exceed VRAM capacity.

Related event: antirez Tunes DwarfStar to 170 t/s Generation on Station(2 posts)→

Original post →

More from Infra

Infra channel →