Tencent open-sources AngelSpec, a speculative decoding framework with up to 2.4x speedup
teortaxesTex · x · 2026-07-29
Tencent has open-sourced AngelSpec, a Torch-native framework for training and deploying speculative decoding drafters.
- It supports both autoregressive MTP drafting and block-parallel DFlash-style decoding.
- On Hy3-A21B, the DFly family reportedly delivers 1.98–2.40× end-to-end speedups over autoregressive decoding across concurrency levels from 4 to 64.
- Tencent also released the training code, Hy3-A21B MTP/DFly drafter weights, a paper, docs, and Hugging Face / ModelScope resources.
- The repo positions AngelSpec as the framework behind the technical report’s released drafters.
Related event: Tencent Open-Sources AngelSpec for Speculative Decoding(2 posts)→
More from Infra
- Colibri-based runner brings Kimi K3 GGUFs to workstation-class hardware — Responsible_Fig_1271 · 2026-07-29
- ML teams still pin torch and Hugging Face deps to keep runs reproducible — vboykis · 2026-07-29
- Jensen Huang Reportedly Committed $4 Billion in Compute to Scale Ilya's AI — iruletheworldmo · 2026-07-29
- AI labs may be hitting compute limits as next leap gets more expensive — haider1 · 2026-07-29
- Kimi K3 pitch says firms spending over $20k a month on APIs may be better off self-hosting — ZeYanjie · 2026-07-29
- Bloom and Seagate ride AI demand while Vertiv’s broader portfolio underperforms — TiernanRayTech · 2026-07-29