NVIDIA Multi-GPU UMAP Processes 870GB of Data in Just 8 Minutes

leland_mcinnes · x · 2026-08-20

NVIDIA cuML and cuVS libraries now support multi-GPU execution for the UMAP algorithm. The new implementation partitions data into balanced clusters, computes local kNN graphs independently, and merges them, avoiding all-to-all communication. Tests on MIRACL and Wiki datasets showed up to a 74x speedup over projected CPU runtimes using eight H100 GPUs, making massive workloads feasible while preserving embedding quality.

Related event: NVIDIA Brings Multi-GPU UMAP, Processing 870GB in 8 Minutes(2 posts)→

Original post →

More from Infra

Infra channel →