Local AI community urges Qwen to bring back a 35B-class MoE for low-VRAM GPUs
julianharris · x · 2026-09-27
A developer is petitioning the Qwen team on behalf of the consumer-GPU local AI community for a new 35B MoE model, arguing Qwen-3.6-35B-A3B is the only model in that range that gave 4GB-16GB VRAM users a real seat in local AI. He also hopes for a Qwen4 n-gram feature and a 9B variant.
More from Models
- Laptop engine streams a 35B model from SSD at 9.4 tok/s, beating GPT-OSS 20B — ImBadGuyInEveryStory · 2026-09-27
- rasbt and marlene_zw break down Claude watermarks, reasoning models in new TechTalk — marlene_zw · 2026-09-27
- Opus 5.5 praised as remarkably efficient: top-tier quality at surprisingly good rates — kimmonismus · 2026-09-27
- Grok 4.7 lifts Terminal-Bench 4.0 from 20.3% to 38%, but burns 125% more output tokens — dl_weekly · 2026-09-27
- TypeSafe's Jev: A Decision-Only Model That Returns Typed Choices Instead of Generated Text — Rahulstark2 · 2026-09-27
- Codex team hints point to speed — GPT-6 Astra on Cerebras rumored — haider1 · 2026-09-27