Running quantized Swift 1.5 in 16GB VRAM at 50 tps: can lower quants help agentic coding?

royalflash417 · reddit · 2026-10-07

A Reddit user reports running a heavily quantized Swift 1.5 model locally on 16GB VRAM + 32GB RAM: the iq2xs quant averages 50 tps (peaks at 73 tps) with 800 t/s prompt processing. They're wondering whether pushing quantization further can reduce KLD and improve agentic coding accuracy, and are asking the community for configs.

Original post →

More from Models

Models channel →