Speculative decoding explains why coding agents type fast on boilerplate but slow on new logic

RunAI_Coder · reddit · 2026-09-09

The author kept blaming "slow afternoons" on the model until reading AMD and Embedded LLM's speculative decoding write-up, which explains the behavior precisely.

How speculative decoding performs

Other clocks in a turn

The author asks whether anyone has logged per-turn TTFT vs decode time vs tool time on a hosted API across a full day, to separate load from one's own cache misses.

Original post →

More from coding & agent

coding & agent channel →