Can a small local model pick the best speculative decoding draft?
Aggravating-Push-207 · reddit · 2026-10-09
A Reddit user proposes a speculative decoding variant: after the target model generates a prefix, multiple candidate drafts are prepared in parallel, and a small, locally-run fast model acts as a judge to pick the best draft instead of relying on a single fixed draft model.
This falls under multi-draft speculative decoding, related to prior work like Medusa and EAGLE.
More from Infra
- Emad Mostaque: OpenAI Burned $10-20M Compute Solving Navier-Stokes, Prices Falling Fast — rohanpaul_ai · 2026-10-09
- Universal Quantum raises $100M+ Series A, largest ever for a UK-based quantum firm — hardimanjames · 2026-10-09
- Texas freezes data center permits as queue balloons from 63 GW to 474 GW in 18 months — elonmusk · 2026-10-09
- Memory stocks are pricing downturns 2-3x deeper than history, Bajarin analysis finds — BenBajarin · 2026-10-09
- Meta's KernelAgent uses multi-agent orchestration for 2.02x Triton kernel speedups — PyTorch · 2026-10-09
- $500/month API bills vs $14k local rig: Mac Studio 512GB or 2x DGX Spark? — rodrigodevbits · 2026-10-09