Qwen 3.8 27b Local Coding Test: 73 tok/s at 128k Context, Usable Performance
julianharris · x · 2026-08-17
A developer tests Qwen 3.8 27b for local coding. With 4-bit unsloth quant on RTX 4090, it achieves 73 tok/s at 128k context, absolutely usable. Disabling thinking mode gives 78 tok/s but much lower quality. Compared to Qwen 3.6 in May, setup is much simpler. Local AI coding agents look promising.
More from coding & agent
- Migrating Hermes Agent from Mac Mini to Hetzner VPS — intellectronica · 2026-08-17
- Study: o3-mini in agentic loop generates high-quality exam questions — mattbeane · 2026-08-17
- Model routing should be optimized at the harness layer, not the gateway — agihouse_org · 2026-08-17
- Opus Ultracodes PCH with Mobile Monitoring — majidmanzarpour · 2026-08-17
- Claude Code Defaults to Auto Mode for Tool Execution — johnmccrea · 2026-08-17
- Open-Source Orca: Run 5 Coding Agents Simultaneously, Compare and Merge Best Result — dr_cintas · 2026-08-17