Prime Flash MoE: Blackwell-Optimized CUDA Kernels Speed Up MoE Inference by 2.4x

pbaylies · x · 2026-08-14

Prime Intellect releases Prime Flash MoE, a set of Blackwell-optimized CUDA kernels for MoE inference. By fusing routing-aware GEMMs, SwiGLU, quantization, and reductions, it avoids materializing intermediate tensors in HBM, saving memory traffic. Up to 2.4x faster than PyTorch grouped GEMM, with 2.3x speedup across 4k-128k tokens. Supports BF16 and MXFP8, with fused and split pipelines. Code is open-sourced.

Related event: Prime Intellect Launches Blackwell-Optimized MoE Inference Kernel(3 posts)→

Original post →

More from Infra

Infra channel →