dotAI talk on llama.cpp speculative decoding: MTP, dflash, dspark
ngxson · x · 2026-09-17
ngxson announced his dotAI conference talk on speculative decoding in llama.cpp, covering MTP, dflash, and dspark, with the replay to be available soon. A useful reference for developers working on local inference optimization.
More from Infra
- Keeping vLLM's prefix cache warm between agent turns — bolts98 · 2026-09-17
- Running Qwen3.8-Flash-Next on 2x5080: should I move from llama.cpp to vLLM? — whatyathinkk · 2026-09-17
- Nvidia's Jensen Huang says chip sales will double next year — zephyr_z9 · 2026-09-17
- GlobalFoundries and Marvell expand Vermont SiGe capacity for AI optical networking — zephyr_z9 · 2026-09-17
- $26,100 desktop AI datacenter: dual RTX PRO 6000 Blackwell workstation goes open source — dee_hw · 2026-09-17
- Huawei Rumored Scale-Up Node with 4,096 Accelerators Could Pack 384TB of HBM — zephyr_z9 · 2026-09-17