Patched vLLM+FlashInfer Pushes Gemma 4 31B to 150 tok/s on a Single B300, Beating SGLang
abhijithneil · x · 2026-09-13
- The author boosted Gemma 4 31B throughput on a single B300 GPU from a 46.7 tok/s eager baseline to 150 tok/s in non-lossy configurations, beating SGLang's best of 139 tok/s
- Key obstacle: Gemma 4's ten global layers use headdim 512, which most attention kernels don't support — FlashInfer in SGLang handled it, but vLLM's path capped at headdim 256
- The author forked vLLM and wrote a patch to lift the headdim 256 limit, and vLLM+FlashInfer hit 150 tok/s vs SGLang+FlashInfer's 139 tok/s
- A blog post with details is coming soon
More from Infra
- Macrocosmos launches IOTA for liquid training on scattered, disaggregated compute — markjeffrey · 2026-09-13
- Dev launches Lorivo: one GPU server serves many LoRA adapters via vLLM — TheOneWhoWil · 2026-09-13
- "The CUDA moat is gone": Japanese neocloud ai& deploys Tenstorrent at scale — DavidBennett__ · 2026-09-13
- MLX vs GGUF on Mac: which local model format and engine wins? — Ok_Warning2146 · 2026-09-13
- The Hugging Bay Brings Torrent Downloads to Open LLM Weights — csuwildcat · 2026-09-13
- Anthropic has lined up compute deals worth up to $517 billion, far above its $180 billion plan — citrini · 2026-09-13