Talk replay: speculative decoding with dflash/dspark speed-ups in llama.cpp

ngxson · x · 2026-10-10

Developer ngxson released the replay of his dotConferences talk, a deep dive into speculative decoding and the dflash/dspark methods, explaining how they improve token generation speed and demonstrating how easy they are to use in llama.cpp. A practical learning resource for local inference optimization.

Related event: dotConf Talk: Speeding Up llama.cpp with Speculative Decoding(2 posts)→

Original post →

More from Infra

Infra channel →