NVIDIA Scales Matrix Factorization to 1M×1M, Doubling Single-GPU Capacity
marc_stampfli · x · 2026-07-30
NVIDIA researchers published a study on accelerated Symmetric Non-negative Matrix Factorization (SymNMF), successfully scaling computations to 1M×1M matrices across 64 GB200 GPUs.
Technical Breakthroughs
- Memory Optimization: A trace-identity reformulation eliminates all $n \times n$ intermediates, roughly doubling single-GPU capacity to handle $n \approx 10^5$.
- Algorithm Evaluation: The team systematically tested 7 algorithm families (over 30 configurations). At the 1M scale, 5 AdaGrad-family methods still successfully converge.
- Performance: Block-SVRG AdaptGrow wins on the ill-conditioned tail-dependence spectrum due to lower per-iteration cost, while full-batch AdaGrad performs best on the dominant-low-rank correlation spectrum.
More from Infra
- CtrlS Building Data Center Capacity Rivaling India's Entire History — shashib · 2026-07-30
- Memory and Chip Stocks Plunge: Kioxia Drops 58%, Samsung Down 37% — CSProfKGD · 2026-07-30
- Building Local AI with Dual AMD R9700s: Is ROCm a Viable Alternative to CUDA? — Syosse-CH · 2026-07-30
- Run Gemma 4 Locally with 16GB RAM: A Zero-Cost Fully Offline Setup Guide — FinanceYF5 · 2026-07-30
- 4090+5060 Ti Hybrid Inference: Runs 122B Model at 37 t/s — Dry_Long3157 · 2026-07-30
- OpenAI to Consume 40% of Global DRAM: The Rise of the Metered Intelligence Complex — Shimano-No-Kyoken · 2026-07-30