NCCL 2.31.2 ships GPU-driven CFT, per-collective tuning and 0-SM collectives for Blackwell-scale training

SkyLi0n · x · 2026-09-23

NVIDIA released NCCL 2.31.2 with notable communication-layer upgrades: GPU-driven Compute Fabric Transport (CFT) host/device APIs for window memory registration and device-side Put/Get/Red/NVLS operations (Blackwell GPUs, CUDA 13.3+), per-collective configuration and tuning (ncclCollConfigt with algorithm selection and CTA/CGA overrides), NVLS+PAT improvements, 0-SM collectives, enhanced multi-NIC/GIN support including an AWS-EFA-contributed GDA backend, and PACE fusion. The release is seen as useful for MoE, FSDP, TP/EP, and Blackwell-scale training.

Related event: NVIDIA Ships NCCL 2.31.2 with GPU-Driven Communication and 0-SM Collectives(2 posts)→

Original post →

More from Infra

Infra channel →