DwarfStar hits 3000 t/s using RAM+VRAM hybrid setup

antirez · x · 2026-08-18

Developer demoed running DeepSeek v4 PRO on DGX Station via DwarfStar orchestration. Despite the model exceeding VRAM, leveraging a peculiar RAM+VRAM hybrid setup achieved a crazy prefill speed of 3000 tokens/s. The author emphasizes needing specific inference engines to exploit hardware for local AI.

Related event: DeepSeek v4 local deployment hits 3000 t/s prefill on DGX Station(2 posts)→

Original post →

More from Infra

Infra channel →