dotConferences talk: how speculative decoding speeds up llama.cpp
ngxson · x · 2026-10-10
The replay of ngxson's talk at dotConferences is now out. He goes deep into speculative decoding and dflash/dspark, explaining how the technique improves token generation speed and showing how easy it is to use with llama.cpp.
Related event: dotConf Talk: Speeding Up llama.cpp with Speculative Decoding(2 posts)→
More from coding & agent
- Building Blender scenes with zero hands-on: an AI-only MCP experiment wins over pros — sidahuj · 2026-10-10
- OpenCode opens more TestFlight slots for its iOS app — DanielLockyer · 2026-10-10
- Anthropic's Claude can now orchestrate up to 1,000 parallel agents via dynamic workflows — The Decoder · 2026-10-10
- AI agent misses $14M in lease clauses: Surge AI's GDP.xlsx benchmarks agents on real spreadsheets — echen · 2026-10-10
- Sentry CEO: coding agent memory means zero context needed to refactor a repo — zeeg · 2026-10-10
- 112 agents, 16 hours, 60,000 lines of code: give agents specs, not prompts — colinmcnamara · 2026-10-10