NVIDIA blog: Gemma 4 hits 10,996 OTSU with Vera Rubin optimizations
ricklamers · x · 2026-08-25
NVIDIA's LPX team published a tech blog detailing recent advances in model inference optimization.
Key Numbers:
- Gemma 4: Achieved 10,996 OTSU speeds (high AL case) and a median of 3,431 OTSU for full 100K input context window workloads.
- 2T-Scale Models: Using Vera Rubin and LPU drafts for best-in-class tok/watt at 1000 tok/s interactivity.
Related event: NVIDIA Showcases Gemma 4 Inference Breaking 10K Speed on LPX(2 posts)→
More from Infra
- Agents need lots of VMs but not lots of CPU/RAM — davidcrawshaw · 2026-08-25
- Developers Rebuild Groq LPU Architecture to Run MicroGPT — IgorCarron · 2026-08-25
- NVIDIA's strategy: investing in open source to counter OpenAI and Anthropic — markjeffrey · 2026-08-25
- Enterprise-managed auth for MCP connectors is now generally available — EricBuess · 2026-08-25
- FreeToken Engine: Run 290B+ Frontier MoE Models Locally on 8GB Laptop GPUs — Saboo_Shubham_ · 2026-08-25
- Meta's AI infrastructure rental won't make it a cloud giant: analyst — DavidLinthicum · 2026-08-25