Debugging slower speeds with MTP enabled on Gemma 4 12B QAT
NovaXeros · reddit · 2026-09-01
A user observed that enabling MTP on Gemma 4 12B QAT reduces inference speed (from 32 to 23 t/s) despite higher VRAM usage on a Radeon 6900XT. The issue persists across different context sizes and model sources. The post seeks help diagnosing the cause and includes detailed launch commands and configuration settings.
More from Infra
- Tencent's Hy4 Preview: 1.25-bit Quantization Cuts Model Size by 7x — ccerrato147 · 2026-09-01
- Opinion: Humanoid robot sales predictions are meant to be forgotten — TiernanRayTech · 2026-09-01
- Hugging Face Transformers Adds Activation Checkpoint Offload to Reduce GPU Memory — StasBekman · 2026-09-01
- Meta Open Sources MetaRoCE Protocol for AI-Scale Ethernet Clusters — Meta_Engineers · 2026-09-01
- Amazon quietly hikes prices on Echo, Kindle, citing rising memory costs — film_girl · 2026-09-01
- Cloudflare turns global network into agentic cloud with Dynamic Workers and AI Gateway — dscape · 2026-09-01