FP8 Tuning Cuts 42.9ms Per Step: Custom SGLang Kernels Boost Inference 126%

HankYeomans · x · 2026-09-23

The author shares first-hand results from optimizing LLM inference on an RTX 6000 Pro.

Original post →

More from Infra

Infra channel →