Qwen3.8 27B DSpark GGUF test: no speedup, high memory usage

Hefty_Wolverine_553 · reddit · 2026-08-15

A Reddit user converts RadixArk/Qwen3.8-27B-DSpark to GGUF and tests it, finding DSpark speculative decoding doesn't improve performance and increases memory usage. On a 5090, 64k context needed for BF16 speculator, 86 t/s, draft acceptance 0.29.

Original post →

More from coding & agent

coding & agent channel →