DwarfStar hits 3000 t/s using RAM+VRAM hybrid setup
antirez · x · 2026-08-18
Developer demoed running DeepSeek v4 PRO on DGX Station via DwarfStar orchestration. Despite the model exceeding VRAM, leveraging a peculiar RAM+VRAM hybrid setup achieved a crazy prefill speed of 3000 tokens/s. The author emphasizes needing specific inference engines to exploit hardware for local AI.
Related event: DeepSeek v4 local deployment hits 3000 t/s prefill on DGX Station(2 posts)→
More from Infra
- MCDMA Enables Direct RDMA Between NVIDIA Spark and Apple Silicon — StephanSturges · 2026-08-19
- Cursor's deep dive on Git at any scale: running Origin like a database — vmg · 2026-08-19
- Gemma 4 Powers Local Enterprise RFP Agent — victormustar · 2026-08-19
- Vector database Qdrant surpasses 250M downloads, recognized by Forrester — qdrant_engine · 2026-08-19
- Mojo programming language is now open source, combining Python ease with C++ performance — mgill25 · 2026-08-19
- Infra Behind Krea 2: Tensor Cores and Crash Philosophy — AI Engineer · 2026-08-19