antirez's DwarfStar: adaptive VRAM/RAM expert placement runs DeepSeek V4 at 45 t/s

antirez · x · 2026-08-17

antirez (Salvatore Sanfilippo, creator of Redis) shares progress on his DwarfStar inference engine for DGX Station: running DeepSeek v4 PRO at Q2 quantization, routed experts are split between VRAM and RAM, with the engine adaptively migrating experts based on past token-generation usage so the hot set stays in VRAM. It currently reaches 45 t/s and can go faster. He calls this just the start of what's possible, reiterating his claim that DwarfStar could be the Station inference engine.

Related event: Redis Author Optimizes DeepSeek V4 to 45 t/s(3 posts)→

Original post →

More from Infra

Infra channel →