Qwen3.8-Flash recipe on DGX Spark: int4 quantization + RDMA, 47.5 t/s code generation
Saren-WTAKO · reddit · 2026-08-31
The author shares a recipe for running Qwen3.8-Flash on a single GB10/DGX Spark, using Intel AutoRound int4 quantization, vLLM, and offloading fp8 ngram table to local SSD or external RDMA server. Detailed performance numbers: 47.5 t/s for code, 60 t/s for JSON, and discusses MTP-induced TTFT overhead.
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- 2 engineers + AI designed a working LLM chip in 2 weeks, no human in the loop — 新智元 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01