ASD: training-free approximate speculative decoding boosts LLM throughput up to 15.26%

新智元 · wechat · 2026-09-05

A joint team from Beihang, Tsinghua, HKU and Peking University proposes Approximate Speculative Decoding (ASD), which relaxes the 'stop at first token mismatch' rule: it selectively accepts low-regret mismatches within a strict request-level budget and reuses the already-computed greedy suffix behind them.

Results:

The paper is explicit that accepting non-greedy tokens changes the trajectory — not lossless — but audits show accuracy holds on most tasks, arguing speed/behavior tradeoffs should be explicitly modeled and auditable.

Original post →

More from Infra

Infra channel →