Windows reset restores RTX 3070 throughput for local Qwen3.6-35B inference

campaigner_ · reddit · 2026-08-04

The author tracked a halved token-per-second issue on an RTX 3070 to a Windows software problem and fixed it by resetting the OS.

After the reset, the setup returned to about 30 tps at an 81,920-token context and could push to 130k context while still staying near 27 tps. The post includes the full llama-server launch command, using Qwen3.6-35B-A3B-UD-Q4KXL.gguf, --gpu-layers 99, --cpu-moe, --cache-type-k/v q80, and other tuning flags.

The main takeaway is that when mysterious throughput regressions persist, a clean Windows reset can be worth trying before spending more time on CUDA and terminal-level debugging.

Related event: Windows Reset Fixes Halved Local LLM Inference Speeds on RTX 3070(2 posts)→

Original post →

More from Infra

Infra channel →