Offload DGX Spark's spec-decode draft model to a spare GPU, freeing memory for context

ciprianveg · reddit · 2026-09-29

The author open-sourced a way for DGX Spark owners to repurpose a spare 10-24 GB GPU: move the speculative-decoding draft model off the Spark onto that GPU over TCP or RDMA — single Spark or cluster setups both supported. The freed GBs can go to extra context or better quant quality. Ships as eugr-vllm compatible mods on GitHub (gb10-vllm/remote-dspark).

Original post →

More from Infra

Infra channel →