CUDA Toolkit 13.4 adds Windows on Arm support, Rubin preview, and finer shared-GPU control

kimmonismus · x · 2026-09-10

NVIDIA's technical blog details CUDA Toolkit 13.4: support for Windows on Arm (extending beyond Linux on Arm), preview functional support for the Rubin architecture (compute capability 107) for early porting, Multi-Process Service V3 with a scriptable CLI, named server instances, TOML config and cgroup-based GPU memory limits for containerized GPU partitioning, plus CUDA Compute Fabric Transport for NVLink communication libraries. CUDA Python gains cuda.core 1.1.0 (texture/surface programming, NUMA-aware managed memory, type stubs) and cuda.compute 1.1 AOT compilation; CCCL 3.4 delivers a faster warp-specialized cub::DeviceScan on Blackwell reaching up to 92% memory-bandwidth utilization. This paves the way for October RTX Spark laptops positioned for local AI models and personal agents.

Related event: CUDA 13.4 Brings Windows on Arm Support Ahead of RTX Spark Laptops(2 posts)→

Original post →

More from Infra

Infra channel →