On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests
dtdisapointingresult · reddit · 2026-09-29
A Reddit user benchmarked ComfyUI's int8 convrot quantized models on DGX Spark expecting 40-50% speedups. Results overturned expectations: Z-Image-Turbo int8 (6.3GB) ran 6.60s vs bf16 (13GB) 7.63s, with nvfp4 fastest at 5.37s — but for H3 5-second video, bf16 (40.2GB) finished in 272.22s, actually beating int8 convrot (21GB) at 286.66s. Dequantization overhead can erase or reverse quantization gains; benchmark on your own hardware before upgrading. Launch flags: --use-ck-attention --disable-mmap --cache-classic.
More from Infra
- SemiAnalysis: Why GLM-5.3 Sparse Attention Doesn't Cut HBM Memory Capacity Needs — burny_tech · 2026-09-29
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B — BAAI · 2026-09-29
- BAAI's MALA attention allocates its own compute, cutting 128K training latency 2.2x — BAAI · 2026-09-29
- Databricks Tops All 4 NVIDIA SOL-ExecBench Kernel Tracks Using AI Agents for ~$70K — Yuchenj_UW · 2026-09-29