Qwen 3.8 27B Local Coding: 73 tok/s at 128K Context on RTX 4090
julianharris · x · 2026-08-17
A developer tests Qwen 3.8 27B locally on RTX 4090 with 4-bit quant and MTP, achieving 73 tok/s at 128K context. Disabling thinking gives 78 tok/s but lower quality. Compared to Qwen 3.6, setup is easier and more stable.
Related event: Qwen 3.8 27B Shows Stable Local Performance(2 posts)→
More from coding & agent
- Migrating Hermes Agent from Mac Mini to Hetzner VPS — intellectronica · 2026-08-17
- Study: o3-mini in agentic loop generates high-quality exam questions — mattbeane · 2026-08-17
- Model routing should be optimized at the harness layer, not the gateway — agihouse_org · 2026-08-17
- Opus Ultracodes PCH with Mobile Monitoring — majidmanzarpour · 2026-08-17
- Claude Code Defaults to Auto Mode for Tool Execution — johnmccrea · 2026-08-17
- Open-Source Orca: Run 5 Coding Agents Simultaneously, Compare and Merge Best Result — dr_cintas · 2026-08-17