A pure Triton W4A16 GEMM claims 1.1x-1.3x decode speedups over cuBLAS FP16
bassrehab · reddit · 2026-07-22
A developer says they wrote a pure Triton W4A16 GEMM kernel that beats cuBLAS FP16 by about 1.1x–1.3x in decode on an A100-SXM4-80GB.
What it does
- Uses FP16 activations with GPTQ/AWQ-style group quantization at group size 128.
- Dequantizes in registers inside the GEMM loop.
- Uses split-K for skinny decode shapes.
Reported results
- qwen2-72b-ffn (8192, 29568): 1.28x faster at M=1, 1.21x at M=8
- llama3-8b-attn-qkv (4096, 4096): 1.33x at M=1, 1.14x at M=8
- llama3-8b-ffn-up (4096, 14336): 1.18x at M=1, 1.10x at M=8
The author notes that cuBLAS still wins once sequences get beyond a few tokens, but argues decode matters more because it runs once per generated token. They also say the code is pure Triton, so it should be structurally portable to AMD, though MI300X validation is still pending.
A writeup and source code are linked, along with a HF Kernel Hub entry.
More from Infra
- AMD and Anthropic expand partnership with up to 2GW of MI450 GPUs and a $5B stake — BenBajarin · 2026-07-22
- Google DeepMind commits $40M in AI tokens and cloud credits for U.S. Energy Department research push — GoogleDeepMind · 2026-07-22
- Jensen Huang says Chinese open models should be used, not feared — hyhieu226 · 2026-07-22
- AWS may be entering a longer brain-drain phase after Dave Brown’s 19-year exit — DavidLinthicum · 2026-07-22
- Glow exits stealth with $180M at a $1.2B valuation to secure AI agents on endpoints — Justgototheeffinmoon · 2026-07-22
- The U.S. Army reportedly blew through a year of AI tokens in roughly one month — Annual_Judge_7272 · 2026-07-22