GPU Collective Latency Optimized to Near Speed-of-Light, Boosting LLM Inference
traviscline · x · 2026-08-25
This arXiv paper investigates optimizing GPU collective communication latency to approach the hardware Speed-of-Light (SoL) limit within scale-up networks. Emerging workloads like long-context LLM inference are increasingly latency-bound. The authors developed low-latency interfaces based on NCCL's device-side API, implementing barrier-free synchronization and efficient symmetric memory usage. Microbenchmarks show substantial latency reductions for small messages, with overhead within 7% of the SoL limit. Integration into real applications improved inter-token latency and throughput for LLM inference and accelerated cuSOLVERMp.
More from Infra
- MLCCs Become Unexpected Bottleneck in AI Racks Amid Cost Surge — tengyanAI · 2026-08-25
- FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution — SteppenAxolotl · 2026-08-25
- West Virginia considers using 50% of data center revenue to cut income tax — bradneuberg · 2026-08-25
- SGLang v0.5.18: Engine 2.38x Faster, AMD NVFP4 Support Added — BanghuaZ · 2026-08-25
- Stoa Intelligence Launches Real-Time Pricing Market for GPU Configs — ycombinator · 2026-08-25
- Neon deep dive: WAL+S3 storage architecture for the era of agents — matei_zaharia · 2026-08-25