8GB RTX 4060 Ti Tuned for 2x-9x Faster Local LLM Inference vs llama.cpp Defaults

ExxploreCraft · reddit · 2026-10-06

The author turned an RTX 4060 Ti (8GB) + 64GB RAM gaming PC into a local inference machine, achieving 2x-9x speedups over default llama.cpp via MoE expert offloading to system RAM, moving display output to the iGPU (+20-30%), native Linux over WSL2 (+33-38%), per-family llama.cpp builds, and KV cache/MTP tuning. Qwen3.6-35B-A3B hit 52-65 tok/s. Open-source repo with one-script install included.

Original post →

More from Infra

Infra channel →