GLM 5.3 flash on dual DGX Sparks gets 50-90% decode boost with new open recipe
swiebertjee · reddit · 2026-10-05
A Redditor reports that a new open NVFP4 recipe (GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark) boosts GLM 5.3 flash decode by 50-90% on dual DGX Sparks:
- Quality benchmarks up 3-13% across B2 categories (prose +13%, bulk SQL INSERT +10%); prefill drops 10-21% as the tradeoff
- Decode now beats DeepSeek v4.0 flash; the newer v4.1 doesn't run on dual Sparks at all
- 40-minute soak: 503 requests, 0 errors, KV pool memory down 72%; the old repetition bug is fixed
- Author ran it for days on coding, devops, and vision tasks and rates it near Claude Opus 4.8 level
The recipe is open on GitHub — a solid empirical reference for local LLM deployment.
More from Infra
- Muse ships reliability fixes after SEVs left cron jobs and scheduled tasks unrecovered — alexandr_wang · 2026-10-05
- AWS Mistakenly Suspends Account, Wabi Down for 3+ Hours With No Recourse — soleio · 2026-10-05
- Is Strix Halo the closest thing to a dream local LLM box? Unified memory vs GPUs for 20B-32B models — Robert__Sinclair · 2026-10-05
- 539 tok/s DeepSeek on 4x RTX 6000 — and a call-out that community benchmarks inflate 20-30% — HankYeomans · 2026-10-05
- Qualcomm's Snapdragon to power next-gen AI assistants for Meta and OpenAI — ryanshrout · 2026-10-05
- Spite: a modular Rust inference engine where every model, GPU and op is pluggable — giveen · 2026-10-05