Debugging slower speeds with MTP enabled on Gemma 4 12B QAT

NovaXeros · reddit · 2026-09-01

A user observed that enabling MTP on Gemma 4 12B QAT reduces inference speed (from 32 to 23 t/s) despite higher VRAM usage on a Radeon 6900XT. The issue persists across different context sizes and model sources. The post seeks help diagnosing the cause and includes detailed launch commands and configuration settings.

Original post →

More from Infra

Infra channel →