NVIDIA's Online Draft Co-Training Speeds Speculative Decoding in Long-Context RL Post-Training
nvidia · hf · 2026-09-09
NVIDIA published Online Draft Co-Training for Speculative Decoding, targeting inference acceleration for large-scale, long-context RL post-training.
Technical highlights:
- Online co-training of draft models to accelerate speculative decoding
- Extends context-parallel attention
- Adds cross-stage feature transport
Addresses generation throughput bottlenecks in RL post-training, especially for long contexts.
More from Infra
- BeaconKV compresses KV cache for long reasoning models via beacon queries — Janghyeon Kim · 2026-09-09
- The 4 things you need to run local AI: models, Hugging Face, runners, and quantization — Roger_M_Taylor · 2026-09-09
- Google says AI servers pay back in under 2 years, just 1 year on its own silicon — SumitGup · 2026-09-09
- Together claims GLM-5.3 Flash beats Claude Fable 5.1 on agentic tasks at ~1% cost — togethercompute · 2026-09-09
- MCP tool failures hide in HTTP 200, agent retries 6 times billed with no error traces — Thirumalaiboobathi · 2026-09-09
- MCP Creator: AI Is Now in a High-Compute Regime, Adjust Your Plans — hrishioa · 2026-09-09