NVIDIA 提出 Draft Co-Training,加速大规模长上下文 RL 后训练推理
nvidia · hf · 2026-09-09
NVIDIA 发布论文 Online Draft Co-Training for Speculative Decoding,针对大规模长上下文 RL 后训练的推理加速。
技术要点:
- 在线协同训练 draft 模型,加速投机解码(speculative decoding)
- 扩展上下文并行注意力(context-parallel attention)
- 引入跨阶段特征传输(cross-stage feature transport)
解决了 RL 后训练中生成吞吐瓶颈,在长上下文场景下尤为重要。
「Infra」频道最新
- BeaconKV:用信标查询预测 KV 缓存复用,压缩长推理链显存 — Janghyeon Kim · 2026-09-09
- 本地跑 AI 只需搞懂 4 件事:模型、Hugging Face、运行器与量化 — Roger_M_Taylor · 2026-09-09
- Google 称 AI 服务器回本周期不到两年,自研芯片仅一年 — SumitGup · 2026-09-09
- Together 称 GLM-5.3 Flash 跑分超 Claude Fable 5.1,成本仅约 1% — togethercompute · 2026-09-09
- MCP 工具失败返回 HTTP 200,Agent 重试 6 次全付费且无告警 — Thirumalaiboobathi · 2026-09-09
- MCP 作者:AI 已进入「高算力机制」,规划要相应调整 — hrishioa · 2026-09-09