Slow TPS on AMD 7900 XT Running Gemma: 131K Context Bottleneck
opoot_ · reddit · 2026-08-13
A user running a Gemma model via LM Studio on Windows reported weirdly slow inference speeds (around 14 TPS). Despite the setup fitting entirely within VRAM using a 131k context length and quantized KV cache, performance suffered significantly.
The user suspects the Windows environment might be a factor and is seeking ideas from the community for performance improvements.
More from Infra
- Menlo Park Daytime Electricity Hits 53.8¢/kWh: Running Own GPUs Becomes Irrational — generativist · 2026-08-13
- Running Krea 2 Turbo on 8GB VRAM: RTX 3070 Ti Local Test — niechta · 2026-08-13
- Fluidstack Visits NYSE to Discuss US AI Infrastructure Investment — MxMnr · 2026-08-13
- Open-Source mlx-dspark Boosts LLM Inference on Mac by 3.3x — A-Rahim · 2026-08-13
- Open-Source CUDA Alternative for Portable AMD GPU Code — tom_doerr · 2026-08-13
- Building Local Open-Weight Agents on a 4GB VRAM GPU — ComplexHuman26 · 2026-08-13