CUDA Toolkit 13.4 adds Windows on Arm support, Rubin preview, and finer shared-GPU control
kimmonismus · x · 2026-09-10
NVIDIA's technical blog details CUDA Toolkit 13.4: support for Windows on Arm (extending beyond Linux on Arm), preview functional support for the Rubin architecture (compute capability 107) for early porting, Multi-Process Service V3 with a scriptable CLI, named server instances, TOML config and cgroup-based GPU memory limits for containerized GPU partitioning, plus CUDA Compute Fabric Transport for NVLink communication libraries. CUDA Python gains cuda.core 1.1.0 (texture/surface programming, NUMA-aware managed memory, type stubs) and cuda.compute 1.1 AOT compilation; CCCL 3.4 delivers a faster warp-specialized cub::DeviceScan on Blackwell reaching up to 92% memory-bandwidth utilization. This paves the way for October RTX Spark laptops positioned for local AI models and personal agents.
Related event: CUDA 13.4 Brings Windows on Arm Support Ahead of RTX Spark Laptops(2 posts)→
More from Infra
- GPU indices bullish, token indices bearish: compute appreciates as intelligence gets commoditized — sudoraohacker · 2026-09-10
- Understanding FlashAttention: A Handbook Tracing FA1 to FA4 and Why HBM Traffic, Not FLOPs, Is the Bottleneck — techNmak · 2026-09-10
- Stealth startup Kepler Computing raises $468M for EUV-free HBM alternative using 3D ferroelectric stacking — Promptmethus · 2026-09-10
- Is multi-tenant cloud computing cooked after the Hugging Face sandbox escape? — WonderFactory · 2026-09-10
- mlx-omarchy 0.4.0: run MLX on the Apple GPU under Linux, no Metal needed — alexcovo_eth · 2026-09-10
- Gemini V4.1 pretraining estimated at ~5e24 FLOPs, two weeks on 4K B300s — teortaxesTex · 2026-09-10