Free 5% speedup: enable CustomAllReduce on SM120 to push tensor parallelism from TP=2 to TP=4

TheZachMueller · x · 2026-10-03

Zach Mueller (Hugging Face) shares a quick inference optimization tip: on SM120 GPUs, instructing your LLM to enable CustomAllRemove—sorry, CustomAllReduce—lets you move from TP=2 up to TP=4 tensor parallelism, netting a free 5% performance boost.

Original post →

More from Infra

Infra channel →