NCCL 2.31.2 ships GPU-driven CFT/RMA, 0-SM collectives, better multi-NIC for Blackwell-scale training
SkyLi0n · x · 2026-09-23
NCCL 2.31.2 is out with GPU-driven CFT/RMA for fine-grained communication, per-collective tuning, NVLS+PAT improvements, 0-SM collectives, improved multi-NIC/GIN support, and PACE fusion. The author sees it as useful for MoE, FSDP, TP/EP, and Blackwell-scale training.
Related event: NVIDIA Ships NCCL 2.31.2 with GPU-Driven Communication and 0-SM Collectives(2 posts)→
More from Infra
- DigitalOcean Managed Agents Enters Public Preview: Pause-When-Idle Cloud Claude Code and Codex — _AustinCalvert_ · 2026-09-23
- ARK analyst: AI is the most powerful joule in history, converting energy to GDP ~10x better than humans — DMaguireARK · 2026-09-23
- MIT's David Bau: have KV caches made RAM permanently expensive? Maybe takeoff is in $/bit — davidbau · 2026-09-23
- AI agents can one-shot kernel exploits, ending containers as a security boundary — OwariDa · 2026-09-23
- AI makes intelligence's materiality visible: reasoning runs on chips, memory and energy — demian_ai · 2026-09-23
- Go.AI Raises $85M Series A for On-Prem AI Infrastructure in Regulated Industries — PTrubey · 2026-09-23