Whittle distills on HF: 27B-A3B MoE quant claimed to run on 8GB VRAM laptops
depressedclassical · reddit · 2026-09-22
A Reddit user highlights logic65 on Hugging Face, who publishes small model distills under the Whittle brand. The 27B-A3B MoE GGUF (based on Qwen3.6-35B) looks like it could run well on a modest 8GB VRAM / 24GB RAM laptop for chat, coding and writing, with a larger Qwen3.8 variant out of reach. Quant quality and speed remain unverified — worth testing locally.
More from Infra
- Meta rumored to partner with Oracle to bring Meta AI Platform to Oracle Cloud users — testingcatalog · 2026-09-22
- Astorias AI launches budget inference service with Qwen at $0.1/M input tokens — me_broke · 2026-09-22
- Osborne, now at OpenAI, says UK datacentre NIMBYs threaten Britain's AI sovereignty — nordicinst · 2026-09-22
- What are KV caches really? A storage expert's explainer from prefill to offload — TheZachMueller · 2026-09-22
- DeepSeek reportedly bets on Huawei chips to train next-gen models; Liang says it 'has to work' — kimmonismus · 2026-09-22
- Modded RTX 3080 20GB laptop GPUs appear on AliExpress for AI use — thatguyjames_uk · 2026-09-22