27B quantized model fixes in 6 min on a 3090 what Gemini Flash couldn't in 40 min
MomentJolly3535 · reddit · 2026-09-28
UkisAI released Swift 1.5 Qwen3.8 27B, a Qwen-based model tuned for token efficiency. A Reddit user benchmarked the IQ4XS quant against Unsloth's Q4KS: UkisAI consistently won at low-thinking, while Unsloth stayed ahead at high-thinking.
The real-world test: a custom Directory Opus script problem stumped Gemini Flash (medium thinking, free Antigravity tier) after 40 minutes of looping and burning the weekly token limit. The same problem was completely fixed by the 27B model in 6 minutes on an old RTX 3090 at 67 t/s. The author didn't expect a quantized 27B in low-thinking mode to beat a major cloud model — calling it a must-have for local inference.
More from Models
- TeleOCR Trends on Hugging Face: A Qwen2.5-VL-Based Chinese Document OCR Model — XingChen-AGI · 2026-09-28
- Kaggle Game Arena: Evaluating LLMs via Head-to-Head Chess, Poker, and Werewolf — kaggle · 2026-09-28
- Perplexity CEO: still using sol 6 for knowledge work — cheap, fast, great compaction — gabriel1 · 2026-09-28
- NerfBench's First Results Find No Nerf: Claude Opus 5.5 Dips Just 0.8% vs Launch — alejandroll10 · 2026-09-28
- Most humans can read this image instantly — most AI vision models can't — JeremyNguyenPhD · 2026-09-28
- AI-generated 7-minute SQLite repo explainer stuns with coherent code walkthrough — deedydas · 2026-09-28