dotConf Talk: Speeding Up llama.cpp with Speculative Decoding

ngxson's dotConferences talk is now available, explaining speculative decoding and how dflash/dspark accelerate llama.cpp generation.

2026-10-10 ~ 2026-10-10 · 2 related posts

1 near-duplicate retellings: ngxson