Same GB300, same workload: switching serving engine moved the benchmark result by 11x

Slight_Republic_4242 · reddit · 2026-09-15

Key finding

SemiAnalysis InferenceX data (Sep 15, partly paywalled) compares Rubin vs GB300, but the more useful part holds hardware fixed and changes only the serving engine:

Takeaway

Quote from the piece: "The engine label is therefore essential when quoting the high-interactivity gain." A throughput number is a pair — hardware plus engine — and a single multiple with no engine named isn't a measurement of anything. The same trap applies to small-scale vLLM/SGLang/llama.cpp deployments.

Original post →

More from Infra

Infra channel →