Slow TPS on AMD 7900 XT Running Gemma: 131K Context Bottleneck

opoot_ · reddit · 2026-08-13

A user running a Gemma model via LM Studio on Windows reported weirdly slow inference speeds (around 14 TPS). Despite the setup fitting entirely within VRAM using a 131k context length and quantized KV cache, performance suffered significantly.

The user suspects the Windows environment might be a factor and is seeking ideas from the community for performance improvements.

Original post →

More from Infra

Infra channel →