NVIDIA's Online Draft Co-Training Speeds Speculative Decoding in Long-Context RL Post-Training

nvidia · hf · 2026-09-09

NVIDIA published Online Draft Co-Training for Speculative Decoding, targeting inference acceleration for large-scale, long-context RL post-training.

Technical highlights:

Addresses generation throughput bottlenecks in RL post-training, especially for long contexts.

Original post →

More from Infra

Infra channel →