FP8 Tuning Cuts 42.9ms Per Step: Custom SGLang Kernels Boost Inference 126%
HankYeomans · x · 2026-09-23
The author shares first-hand results from optimizing LLM inference on an RTX 6000 Pro.
- FP8 tuning was the biggest win: stock configs were tuned for H100 defaults; removing them and adding RTX 6000 Pro-specific settings for W8A8 FP8 GEMM decoding shaved 42.9ms off a 70ms baseline step (+126% improvement).
- How: they forked SGLang kernels and added configuration specific to the RTX 6000 Pro. Kernel fallback ran at only 80GB/s versus 750GB/s after tuning (on a 1600GB/s GPU).
- Takeaway: shaving milliseconds per step in decoding can yield the largest overall gains in this kind of experiment.
More from Infra
- Cloudflare Bets on Microtransactions to Save the Web, but Agents Want Answers, Not Fragments — Ronangmi · 2026-09-23
- Google Colab joins Google AI plans with faster GPUs and background execution — DynamicWebPaige · 2026-09-23
- NEAR AI confidential inference goes live on Bittensor via SayGm's TDX-enclave routing — markjeffrey · 2026-09-23
- 6 serving-side techniques that make LLM inference faster - from prefix caching to PD disaggregation — techNmak · 2026-09-23
- SanDisk shares surge 7.56% as Wall Street turns bullish on AI memory demand — Polymarket · 2026-09-23
- SGLang v0.5.20 ships: Intel XPU support, up to 52% faster decode — hsu_byron · 2026-09-23