Gemma 4 MTP Slows Down on Dual 3090s

DjCanalex · reddit · 2026-07-19

The author tested the MTP (draft/multi-token prediction) performance of Gemma 4 31B QAT on dual and single RTX 3090 setups, concluding that it actually slows down in a dual-GPU environment.

Observed Results

Additional Logs

The author shared output logs for both configurations along with their llama-server startup commands, seeking to understand why MTP performance on multiple GPUs contradicts the single-GPU results.

Original post →

More from Infra

Infra channel →