Windows reset restores RTX 3070 throughput for local Qwen3.6-35B inference
campaigner_ · reddit · 2026-08-04
The author tracked a halved token-per-second issue on an RTX 3070 to a Windows software problem and fixed it by resetting the OS.
After the reset, the setup returned to about 30 tps at an 81,920-token context and could push to 130k context while still staying near 27 tps. The post includes the full llama-server launch command, using Qwen3.6-35B-A3B-UD-Q4KXL.gguf, --gpu-layers 99, --cpu-moe, --cache-type-k/v q80, and other tuning flags.
The main takeaway is that when mysterious throughput regressions persist, a clean Windows reset can be worth trying before spending more time on CUDA and terminal-level debugging.
Related event: Windows Reset Fixes Halved Local LLM Inference Speeds on RTX 3070(2 posts)→
More from Infra
- Nuclear Startup Valar Raises $1B Led by Sequoia to Scale Reactors — kleffew94 · 2026-08-04
- AI API revenue still trails hyperscaler capex by a wide margin in 2025 chart — SurpriseDog9000 · 2026-08-04
- U.S. heartland backlash grows as AI data centers reshape local communities — altryne · 2026-08-04
- Semiconductors and data centers are being built far slower than AI demand — robleclerc · 2026-08-04
- Gemma 4 31B can use over 13× more KV-cache memory than DeepSeek V4 Flash — teortaxesTex · 2026-08-04
- MCP server brings structured compile, flash and stateful GDB to embedded boards — Historical_Court795 · 2026-08-04