TorchSpec open-sources 3 SOTA speculative decoding draft models for Kimi K3
hongyangzh · x · 2026-09-04
The TorchSpec team has released the Kimi K3 Draft Collection: three SOTA speculative decoding draft models (EAGLE-3, DFlash2, DSpark) trained with vLLM on NVIDIA GB200, with data recipes published. They also note a key empirical finding: high acceptance rate doesn't equal faster decoding — each draft architecture has its own drawbacks.
Related event: Kimi K3 gets open-sourced speculative decoding draft models(2 posts)→
More from Infra
- HPE Delivers Strong Q3 on AI Server Demand, Raises FY2026 Outlook — mattwbaker · 2026-09-04
- All Chromium Browsers Hit by Hard-to-Reproduce Data Loss Bug, Devs Say — uwukko · 2026-09-04
- Hundreds protest Scotland's datacentre boom, demanding pause on 20+ proposed projects — nordicinst · 2026-09-04
- OpenAI, Anthropic and xAI hit by simultaneous outages, disrupting all three leading AI providers — pstAsiatech · 2026-09-04
- Built a Dual RTX 6000 Pro Rig for Local DeepSeek — Warns Against Influencer Build Hype — HankYeomans · 2026-09-04
- Inference engines are an underexamined attack surface, self-hosting ops warned — JeremyCMorgan · 2026-09-04