Ds4 v0.6.2: Runs DeepSeek 284B on single DGX Spark at 1000 tok/s
pbaylies · x · 2026-08-19
Ds4 v0.6.2 introduces a real memory budgeting mechanism that measures actual usage instead of static reservation, enabling robust scaling in shared environments. Benchmarks show it can stably serve the 284B parameter DeepSeek V4 Flash model on a single DGX Spark at 1000 tok/s prefill and 59 tok/s for multi-agent serving. The design implements demand-mapped context and graceful rejection upon resource exhaustion instead of OOM crashes.
More from Infra
- Marvell grants Google warrant as part of expanded custom AI chip deal — firstadopter · 2026-08-19
- SALT: CELF-Based Sentence-Level Compression for KV Cache Retrieval — No_Sky9786 · 2026-08-19
- Beyond human intuition: AI designs chip components 500x smaller than engineering limits — ChuckDBrooks · 2026-08-19
- ComfyUI becomes unusable overnight with MiniMax H3, causing system freezes — Fit-Association-448 · 2026-08-19
- Suggestion: Move Anthropic bio AI to Tenstorrent to cut costs — DavidBennett__ · 2026-08-19
- Palantir-Powered Sovereign AI Accelerates Autonomy and Ops — CeoOndas · 2026-08-19