Speculative decoding on or off: the 35B MoE offloading question on 8GB VRAM

Infinite-Local5435 · reddit · 2026-09-14

A Reddit user asks whether to enable speculative decoding in CPU/GPU offloading scenarios, running a 35B A3B MoE model on an 8GB RTX 5060 with 32GB RAM via llama-server.

Current setup: 40 tok/s generation and 500 tok/s prompt processing with no MTP, flash attention on, q8 KV cache, q4KXL quantization, 4096 batch/ubatch, 16 CPU cores plus MoE offload.

They recall a community consensus from a few months ago favoring NTP for speed, but with reports it heavily slowed prompt processing, and ask whether that still holds and how others configure llama-server for speculative decoding.

Original post →

More from Infra

Infra channel →