Tencent Hunyuan open-sources AngelSpec, a full-stack speculative decoding framework
腾讯混元 · wechat · 2026-07-29
Tencent Hunyuan open-sourced AngelSpec, an end-to-end speculative decoding framework that spans drafter training, architecture design, and online deployment.
- It also released Hy3-A21B MTP and DFlydrafter weights plus training code, making the whole stack reproducible.
- The new DFly drafter reports better acceptance length and throughput than DFlash and MTP, with average speedups of 1.98–2.40× over the AR baseline across 4–64 concurrency, peaking at 2.86× on code/math.
- The D-cut serving policy adds another +15.7% throughput under real high-concurrency traffic by treating verification as a shared batch resource.
- For dialogue, MTP+TTT raises average acceptance rate from 52.8% to 66.4% and acceptance length from 2.58 to 2.99 by fixing train-inference mismatch.
- The framework itself is built on TorchSpec, supports 128k long-context training, and includes real serving-side acceptance evaluation.
Related event: Tencent Open-Sources AngelSpec for Speculative Decoding(2 posts)→
More from Infra
- Tencent open-sources AngelSpec, a speculative decoding framework with up to 2.4x speedup — teortaxesTex · 2026-07-29
- Kimi K3 pitch says firms spending over $20k a month on APIs may be better off self-hosting — ZeYanjie · 2026-07-29
- Bloom and Seagate ride AI demand while Vertiv’s broader portfolio underperforms — TiernanRayTech · 2026-07-29
- llama.cpp merges speculative decoding support for GLM-5.2 — YPSONDESIGN · 2026-07-29
- ChipAgents raises $60M Series A2 and says ARR is up 6x in H1 2026 — WilliamWangNLP · 2026-07-29
- NeurIPS 2026 workshop seeks papers on decentralized foundation model training — peter_richtarik · 2026-07-29