27B quantized model fixes in 6 min on a 3090 what Gemini Flash couldn't in 40 min

MomentJolly3535 · reddit · 2026-09-28

UkisAI released Swift 1.5 Qwen3.8 27B, a Qwen-based model tuned for token efficiency. A Reddit user benchmarked the IQ4XS quant against Unsloth's Q4KS: UkisAI consistently won at low-thinking, while Unsloth stayed ahead at high-thinking.

The real-world test: a custom Directory Opus script problem stumped Gemini Flash (medium thinking, free Antigravity tier) after 40 minutes of looping and burning the weekly token limit. The same problem was completely fixed by the 27B model in 6 minutes on an old RTX 3090 at 67 t/s. The author didn't expect a quantized 27B in low-thinking mode to beat a major cloud model — calling it a must-have for local inference.

Original post →

More from Models

Models channel →