Prime Intellect Releases Blackwell-Optimized MoE Inference Kernels with Fused Routing and Quantization

afurgs · x · 2026-08-14

Prime Intellect introduces Prime Flash MoE, a set of CUDA kernels optimized for Blackwell GPUs, fusing routing-aware GEMMs, SwiGLU, quantization, and reductions without materializing intermediates. It supports BF16 and MXFP8 paths and is benchmarked on B200s.

Related event: Prime Intellect Launches Blackwell-Optimized MoE Inference Kernel(3 posts)→

Original post →

More from Infra

Infra channel →