dotConferences talk: how speculative decoding speeds up llama.cpp

ngxson · x · 2026-10-10

The replay of ngxson's talk at dotConferences is now out. He goes deep into speculative decoding and dflash/dspark, explaining how the technique improves token generation speed and showing how easy it is to use with llama.cpp.

Related event: dotConf Talk: Speeding Up llama.cpp with Speculative Decoding(2 posts)→

Original post →

More from coding & agent

coding & agent channel →