Prime Flash MoE: Blackwell-Optimized CUDA Kernels Speed Up MoE Inference by 2.4x
pbaylies · x · 2026-08-14
Prime Intellect releases Prime Flash MoE, a set of Blackwell-optimized CUDA kernels for MoE inference. By fusing routing-aware GEMMs, SwiGLU, quantization, and reductions, it avoids materializing intermediate tensors in HBM, saving memory traffic. Up to 2.4x faster than PyTorch grouped GEMM, with 2.3x speedup across 4k-128k tokens. Supports BF16 and MXFP8, with fused and split pipelines. Code is open-sourced.
Related event: Prime Intellect Launches Blackwell-Optimized MoE Inference Kernel(3 posts)→
More from Infra
- SanDisk KV Cache TAM Estimate Questioned: May Be Inflated by 2x-4x — zephyr_z9 · 2026-08-14
- AMAT earnings: vague growth guidance, waiting for Intel visibility — BenBajarin · 2026-08-14
- ComfyUI ROCm Performance Doubled: CK Attention and Dynamic VRAM Benchmarks — VQSGecko · 2026-08-14
- Unsloth releases Qwen3.8-27B quantized: NVFP4 1.5x faster, retains 92-97% accuracy — danielhanchen · 2026-08-14
- Cloudflare Workers Now Support Python, Run on Global Network — craigsdennis · 2026-08-14
- Unsloth Releases GGUF Quantized Files for Qwen3.8-27B for Local Deployment — apitman · 2026-08-14