Prime Flash MoE: Blackwell-Optimized CUDA Kernels for MoE Inference
sloppenheimer · x · 2026-08-14
PrimeIntellect introduced Prime Flash MoE, a set of Blackwell-optimized CUDA kernels designed for Mixture-of-Experts (MoE) inference.
Core Optimizations & Performance:
- Avoids materializing intermediate tensors and activations (in fused mode), significantly saving memory traffic.
- Delivers up to 2.4× speedup over PyTorch grouped GEMM and 2.3× speedup across the 4k-128k token range.
- Integrated into their prime-rl framework to accelerate the forward pass of MoE models.
Dual Data Paths:
- Features paths for both BF16 and MXFP8, sharing the same structure but differing in how they feed tensor cores.
- Offers fused and split pipelines. Fusion avoids activation round-trips for small problem sizes, while the split pipeline becomes more cost-effective as the working set grows by materializing smaller activations.
More from Infra
- 4x Faster Kimi-K3 Decoding: Red Hat Releases DSpark Speculator — woosuk_k · 2026-08-14
- W&B Crosses 1 Billion AI Training Runs After 9 Years — morgymcg · 2026-08-14
- Nvidia Partners with Finance Giants to Mobilize $500B, Turning AI Compute into an Asset Class — sanjaykalra · 2026-08-14
- NVIDIA's NeMo Switchyard: Model Routing as the Agent Budget Manager — krishnan · 2026-08-14
- Nvidia and Wall Street Aim to Mobilize $500B: Is Compute Really an Asset Class? — whurley · 2026-08-14
- Running MiniMax H3 on RTX 5060 Ti: Resolution is the Real Bottleneck — danielcar · 2026-08-14