vLLM v0.30 perf: 3.23x Gemma 4 multimodal speedup, ROCm TPOT -26.4%
vllm_project · x · 2026-09-23
Performance highlights from vLLM v0.30.0's release thread:
- RL sampling masks: GPU compaction restores throughput to within 2–3% of no-mask at 512 requests, fixing an 2x step-time regression
- Kimi K3 mixed batches: 5.2–7.7% higher E2E throughput
- Gemma 4 multimodal prefixes: 3.23x E2E speedup on RTX PRO 6000
- AMD ROCm: packed W4A16 zero-points cut median TPOT 26.4% on Gemma 4 AWQ
- Arm CPU: 6% higher throughput on Llama 3.1 8B W8A8
Related event: vLLM v0.30.0 Released with 762 Commits and Major Performance Gains(3 posts)→
More from Infra
- tinygrad hits ~200 tok/s MiMo-V2.6-Pro on MI300X, brought up via GLM-5.3 — AIFlow_ML · 2026-09-23
- Rumors: 64GB+ VRAM RTX 5090 in R&D but not coming anytime soon — AIFlow_ML · 2026-09-23
- Swarm-built inference engine runs Qwen Image-2.1: 1K images in under 0.5s — bingxu_ · 2026-09-23
- MLX MTP head silently ignored: a 3-line fix boosts Mac local decode speed by up to 79% — Micha0827 · 2026-09-23
- Pluton: open-source self-hosted backup platform wrapping Restic and Rclone for encrypted cloud replication — tom_doerr · 2026-09-23
- TensorSharp's logit-reading approach beats LocalJev at structured decisions, 3.3x faster — fuzhongkai · 2026-09-23