Talk replay: speculative decoding with dflash/dspark speed-ups in llama.cpp
ngxson · x · 2026-10-10
Developer ngxson released the replay of his dotConferences talk, a deep dive into speculative decoding and the dflash/dspark methods, explaining how they improve token generation speed and demonstrating how easy they are to use in llama.cpp. A practical learning resource for local inference optimization.
Related event: dotConf Talk: Speeding Up llama.cpp with Speculative Decoding(2 posts)→
More from Infra
- TokenSpeed hits SOTA for Kimi K3 on AMD MI355X: 1.4x faster than ATOM at 216 tok/s — zhyncs42 · 2026-10-10
- SemiAnalysis: NVIDIA Rubin with vLLM Delivers 3.2x Profit per Gigawatt, up to 10x Perf per Dollar vs GB300 NVL72 — woosuk_k · 2026-10-10
- Nebius Spot Pricing Debut: H200 Preemptible VMs Diverge from $0.79 in EU to $5.30 in us-central1 — kevinsxu · 2026-10-10
- Data firms need frontier-level training infra to prove data value, says Proximal — AccBalanced · 2026-10-10
- Cloudflare ships open-weight Clef-omni with audio/video input, cuts Clef-flash to $0.038/M tokens — Cloudflare Blog · 2026-10-10
- Musk says Grokbot will route to best models like Claude, betting models commoditize — JOBhakdi · 2026-10-10