dotConf Talk: Speeding Up llama.cpp with Speculative Decoding
ngxson's dotConferences talk is now available, explaining speculative decoding and how dflash/dspark accelerate llama.cpp generation.
2026-10-10 ~ 2026-10-10 · 2 related posts
- Talk replay: speculative decoding with dflash/dspark speed-ups in llama.cpp — ngxson · 2026-10-10
1 near-duplicate retellings: ngxson