FlashInfer Merges Blackwell Kernels, Boosting E2E Inference Speed by 7.6%
YouJiacheng · x · 2026-08-06
FlashInfer has officially merged the tinygemm2 kernels optimized for the Blackwell architecture (SM100/SM103).
- Performance: Delivers up to 1.79x faster speeds on large shapes, with an 18-23% geomean reduction in kernel time.
- End-to-End Speedup: Translates to up to a 7.6% end-to-end serving performance boost in SGLang.
- Compatibility: Ensures bit-identical outputs compared to the reference implementation and gracefully falls back on unsupported architectures.
More from Infra
- Goldman Sachs: Hyperscaler Capex to Hit $1.2 Trillion by 2027 — Beth_Kindig · 2026-08-06
- Cloudflare Unveils Kitesurf, a Browser Built for AI Agents — threepointone · 2026-08-06
- Cloudflare Introduces Fully Stateless Next-Gen MCP Architecture — threepointone · 2026-08-06
- Cloudflare AI Search Update: Build a Dedicated Search Engine for Your Agents — Cloudflare Blog · 2026-08-06
- slime v0.3.1 Released: Major Optimizations for Memory and Training Performance — hsu_byron · 2026-08-06
- Making the Rust Compiler 3x Faster Could Save Hundreds of Millions Annually — doodlestein · 2026-08-06