New theory extends speculative decoding acceptance beyond distribution-preserving sampling
baseten · x · 2026-10-01
Aaryam Sharma builds a theory for when speculative decoding draft tokens get accepted under non-standard regimes. Standard speculative sampling has an elegant formula: acceptance probability is 1 − TV(p, q). But production systems also run greedy decoding, relaxed acceptance, and tree-based acceptance, which previously had only conditions (argmax q = argmax p) rather than exact formulas. The thread develops unified theory for these cases.
More from Infra
- Pareto partners with Engy to bring Bittensor SN10 inference optimization to external customers — markjeffrey · 2026-10-01
- Respan launches Span-01 router: 37% cheaper than best single model at matching accuracy — ycombinator · 2026-10-01
- RTX 3090 power can be dialed down to 120W for inference, saving power and heat — QuixiAI · 2026-10-01
- AI as compiler: model writes PTX directly, 1.37x speedup on FlashAttention over Triton — Azaliamirh · 2026-10-01
- Dell ships first Vera Rubin NVL72 rack-scale systems in volume, citing unprecedented NVIDIA partnership — yenkel · 2026-10-01
- Shibaura Tech's Ozaki Scheme II lands in CUDA 13.4, squeezing FP64 from AI-focused GPUs — udmrzn · 2026-10-01