Running quantized Swift 1.5 in 16GB VRAM at 50 tps: can lower quants help agentic coding?
royalflash417 · reddit · 2026-10-07
A Reddit user reports running a heavily quantized Swift 1.5 model locally on 16GB VRAM + 32GB RAM: the iq2xs quant averages 50 tps (peaks at 73 tps) with 800 t/s prompt processing. They're wondering whether pushing quantization further can reduce KLD and improve agentic coding accuracy, and are asking the community for configs.
More from Models
- OpenAI reasoning models went from basic arithmetic to decades-old math breakthroughs in two years — daniel_mac8 · 2026-10-07
- Gary Marcus: OpenAI's Math Proofs Rely on Symbolic AI, Vindicating His Stance but Not AGI — GaryMarcus · 2026-10-07
- Ex-Googler suspects Project Astra voice got quantized and downgraded — joannejang · 2026-10-07
- OpenAI Reportedly Makes Serious Breakdown Toward Riemann Hypothesis — ipeirotis · 2026-10-07
- Jev-as-a-Judge: New Model Boosts LLM Judge Reliability for Agent Evals — omarsar0 · 2026-10-07
- X rolls out @bot tagging: reply to any post to save to Notion, set reminders via Grok — tetsuoai · 2026-10-07