PyTorch 新文档揭 Grace 架构 .cpu() 隐性瓶颈:复用 pinned 缓冲区

StasBekman · x · 2026-10-07

NVIDIA 的 Krishna Kalyan 正在为 PyTorch 提交文档 PR,讲清 NVLink-C2C 系统(Grace Hopper、Grace Blackwell)上 CUDA 到 CPU 传输的性能细节,合并前公开征询意见。

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →