Free 5% speedup: enable CustomAllReduce on SM120 to push tensor parallelism from TP=2 to TP=4
TheZachMueller · x · 2026-10-03
Zach Mueller (Hugging Face) shares a quick inference optimization tip: on SM120 GPUs, instructing your LLM to enable CustomAllRemove—sorry, CustomAllReduce—lets you move from TP=2 up to TP=4 tensor parallelism, netting a free 5% performance boost.
More from Infra
- ERRC poster: entropy-reinvested residual correction for tensor-parallel LLM inference — PyTorch · 2026-10-03
- Dev platform cuts prices in half, making GitHub Action runners 12x cheaper — aniketmaurya · 2026-10-03
- Hand-written Blackwell GEMM in CuTe DSL hits 1401 TFLOP/s, about 97% of cuBLAS on B200 — retr0jirachi · 2026-10-03
- David Patterson: solar's near-vertical cost curve will power the singularity — davidpattersonx · 2026-10-03
- Local LLM setup: Strata lets a 7900XTX + 64GB RAM run Qwen Flash at 60 tok/s at 250K context — soyalemujica · 2026-10-03
- DGX Spark shortage derails $8,800 donation plan as buyer can't find stock — cyrus_zei · 2026-10-03