Hot-swappable speculative decoding boosts Qwen 27B on 16GB CUDA

tsangberg · reddit · 2026-09-13

Reddit user tsangberg built on Raymond's KV cache streaming fork of llama.cpp to add hot-swappable speculative decoding, open-sourced as llama.cpp-adaptive-kv-streaming.

Original post →

More from coding & agent

coding & agent channel →