PyTorch PR documents hidden .cpu() bottleneck on Grace NVLink-C2C systems
StasBekman · x · 2026-10-07
A PyTorch docs PR by Krishna Kalyan (NVIDIA) documents CUDA-to-CPU transfers on NVLink-C2C systems (Grace Hopper, Grace Blackwell), open for review before merging.
- On C2C systems, CPU↔GPU copies automatically use C2C, but the bottleneck is the CPU side: tensor.cpu() allocates fresh pageable memory, so the first copy is limited by page allocation and faults, not link bandwidth.
- Activation offloading that calls .cpu() every step pays this cost repeatedly.
- The fix: allocate pinned CPU buffers once and reuse them; the docs also cover synchronization rules, NUMA placement, and the caveat that pinned memory can still fault under OS memory compaction.
More from Infra
- Streaming MoE Experts from SSD: DeepSeek V4.1 Hits ~40 tps via mlx-stream — HankYeomans · 2026-10-08
- Cloudflare's $200B machine-payments play: 75M robot payments at $0.30 a time — LexSokolin · 2026-10-08
- GPU rental prices rise across all generations: Nebius hikes H100 by 17%, B300 by 21% — Beth_Kindig · 2026-10-08
- TypeSafeAI served a trillion tokens on Modal within four days of Jev launch — josh_wills · 2026-10-08
- OpenAI Compute 'Speedrun': Then-vs-Now Comparison Shows Explosive Scale-Up — ns123abc · 2026-10-08
- The endgame AI: a tiny local model across your phone, laptop and glasses that knows you best — VraserX · 2026-10-08