Qwen3.8 27B DSpark GGUF test: no speedup, high memory usage
Hefty_Wolverine_553 · reddit · 2026-08-15
A Reddit user converts RadixArk/Qwen3.8-27B-DSpark to GGUF and tests it, finding DSpark speculative decoding doesn't improve performance and increases memory usage. On a 5090, 64k context needed for BF16 speculator, 86 t/s, draft acceptance 0.29.
More from coding & agent
- Using Agent Trajectories to Mine Data for Model Distillation and Eval — hwchase17 · 2026-08-15
- Developer uninstalls AI agent Hermes after complex setup process — tristanbob · 2026-08-15
- Dev Accidentally Burns $1,475 in API Credits via Fast Mode Bug — bytebot · 2026-08-15
- DSH vs Pi: Extreme Plugin Architecture and Immutable Session Event Streams — dotey · 2026-08-15
- Anthropic Architect: You Need Graph Engineering, Not Just Prompts — goyalshaliniuk · 2026-08-15
- Opinion: Real LLM engineering involves evals and versioning, not just prompting — bgoncalves · 2026-08-15