Can a small local model pick the best speculative decoding draft?

Aggravating-Push-207 · reddit · 2026-10-09

A Reddit user proposes a speculative decoding variant: after the target model generates a prefix, multiple candidate drafts are prepared in parallel, and a small, locally-run fast model acts as a judge to pick the best draft instead of relying on a single fixed draft model.

This falls under multi-draft speculative decoding, related to prior work like Medusa and EAGLE.

Original post →

More from Infra

Infra channel →