Same GB300, same workload: switching serving engine moved the benchmark result by 11x
Slight_Republic_4242 · reddit · 2026-09-15
Key finding
SemiAnalysis InferenceX data (Sep 15, partly paywalled) compares Rubin vs GB300, but the more useful part holds hardware fixed and changes only the serving engine:
- At a 170 TPS interactivity target, Rubin shows 62.9x the tokens per megawatt of GB300 running TRTLLM; switch GB300 to SGLang and Rubin's lead drops to 5.56x — the GB300 result moved 11x on engine choice alone
- Caveat: 170 TPS sits near the fastest point of TRTLLM's measured curve, so it's TRTLLM falling off its own curve, not SGLang being magic; at 100 TPS the two are much closer (21.1M vs 28.5M tok/s/MW)
- Ordering flips by target: TRTLLM leads at 75 TPS, SGLang leads at higher targets
- GB300 on SGLang reaches roughly the same max interactivity as Rubin — buried in an article titled as a hardware multiple
Takeaway
Quote from the piece: "The engine label is therefore essential when quoting the high-interactivity gain." A throughput number is a pair — hardware plus engine — and a single multiple with no engine named isn't a measurement of anything. The same trap applies to small-scale vLLM/SGLang/llama.cpp deployments.
More from Infra
- Orthrus Study: Lossless Speculative Decoding Holds Only at High Numerical Precision — Ilya Koziev · 2026-09-15
- DeepSeek V4.1 Flash on M3 Ultra nearly doubles decode to 31 t/s with first public DSpark Metal port — IngeniousIdiocy · 2026-09-15
- Trigger.dev's chat.agent turns AI chats into durable tasks that survive crashes and redeploys — CodeByPoonam · 2026-09-15
- Oracle Executes Pre-Dawn Mass Layoffs as AI Data Center Spending Balloons — 量子位 · 2026-09-15
- Solar panels now ~$0.12/watt, down from $5-6, seen as key to data center growth — kimmonismus · 2026-09-15
- Looking for the fastest local vision model to run on an RTX 5090 — StartupTim · 2026-09-15