Speculative decoding on or off: the 35B MoE offloading question on 8GB VRAM
Infinite-Local5435 · reddit · 2026-09-14
A Reddit user asks whether to enable speculative decoding in CPU/GPU offloading scenarios, running a 35B A3B MoE model on an 8GB RTX 5060 with 32GB RAM via llama-server.
Current setup: 40 tok/s generation and 500 tok/s prompt processing with no MTP, flash attention on, q8 KV cache, q4KXL quantization, 4096 batch/ubatch, 16 CPU cores plus MoE offload.
They recall a community consensus from a few months ago favoring NTP for speed, but with reports it heavily slowed prompt processing, and ask whether that still holds and how others configure llama-server for speculative decoding.
More from Infra
- 15 self-funded GPU jobs show cost estimators overshoot by a third — Worldly_North_7213 · 2026-09-14
- Kokoro TTS ported to Apple Core AI: 54 voices running fully on-device with zero API cost — amos_gyamfi · 2026-09-14
- Running an LLM agent on a 512MB board with decoupled memory and live cross-machine migration — D777Castle · 2026-09-14
- Across 3,171 sessions and 30B tokens, only 0.3% was model output — a local tool that searches your agent logs — Rare_Guide_9830 · 2026-09-14
- Maia 200 hits ~12 TFLOP/s FP4 in 1mm²: density should be a first-class goal — thoefler · 2026-09-14
- Dev's 24/7 self-hosted AI stack: OpenWebUI, pidot, Tailscale, GLM and DeepSeek — andfanilo · 2026-09-14