NCCL 2.31.2 ships GPU-driven CFT, per-collective tuning and 0-SM collectives for Blackwell-scale training
SkyLi0n · x · 2026-09-23
NVIDIA released NCCL 2.31.2 with notable communication-layer upgrades: GPU-driven Compute Fabric Transport (CFT) host/device APIs for window memory registration and device-side Put/Get/Red/NVLS operations (Blackwell GPUs, CUDA 13.3+), per-collective configuration and tuning (ncclCollConfigt with algorithm selection and CTA/CGA overrides), NVLS+PAT improvements, 0-SM collectives, enhanced multi-NIC/GIN support including an AWS-EFA-contributed GDA backend, and PACE fusion. The release is seen as useful for MoE, FSDP, TP/EP, and Blackwell-scale training.
Related event: NVIDIA Ships NCCL 2.31.2 with GPU-Driven Communication and 0-SM Collectives(2 posts)→
More from Infra
- ARK analyst: AI is the most powerful joule in history, converting energy to GDP ~10x better than humans — DMaguireARK · 2026-09-23
- MIT's David Bau: have KV caches made RAM permanently expensive? Maybe takeoff is in $/bit — davidbau · 2026-09-23
- AI agents can one-shot kernel exploits, ending containers as a security boundary — OwariDa · 2026-09-23
- AI makes intelligence's materiality visible: reasoning runs on chips, memory and energy — demian_ai · 2026-09-23
- Go.AI Raises $85M Series A for On-Prem AI Infrastructure in Regulated Industries — PTrubey · 2026-09-23
- InstinctFlash: open-source runtime runs 8 robot model families in real time on one commercial GPU — chris_j_paxton · 2026-09-23