AWS benchmarks G7 Blackwell vs G5/G6 for 30B MoE inference on SageMaker
AWS ML Blog · rss · 2026-09-09
AWS ML Blog benchmarked two 30B-class MoE models across three GPU instance families on SageMaker AI, showing gains from the new G7 instances powered by RTX PRO 4500 Blackwell GPUs.
- Use case 1: Qwen3-Coder-30B (FP8) compared apples-to-apples on ml.g5/g6/g7.12xlarge with the DJL LMI 28.0 container — G5/G6 use four GPUs (96GB total) while G7 uses two (64GB).
- Use case 2: NVIDIA Nemotron-3-Nano-30B-A3B-NVFP4 evaluated across G6/G6e/G7 via the Generative AI Inference Recommendations feature with vLLM.
- Key takeaways: MoE decoding is memory-bandwidth bound, and NVFP4 4-bit quantization gets native hardware acceleration only on G7's Blackwell Tensor Cores, giving G7 a structural advantage.
- G7 is currently GA in US East (Ohio) and US West (Oregon) only.
More from Infra
- Oligopoly Equilibrium: why semiconductor markets settle at ~3 players — BenBajarin · 2026-09-09
- Deft Robotics launches unified deployment platform for physical AI — Scobleizer · 2026-09-09
- One user's local AI rig: DGX Spark bandwidth disappoints, RTX Pro 6000 looks like a steal — EAccelerate_42 · 2026-09-09
- Perplexity CEO: inference now served on NVLink Blackwells, Vera Rubin next — AravSrinivas · 2026-09-09
- TSMC to start High-NA EUV production in 2030 as ASML gains broader supply roadmap — BenBajarin · 2026-09-09
- Companies pay 10-20x more for cloud AI infrastructure than on-prem, analyst warns — DavidLinthicum · 2026-09-09