Qwen 3.8 27B Vision MLX Quants: Low-Bit Performance Benchmarks
Top-Eye-8104 · reddit · 2026-08-26
The team released various MLX quantizations of the Qwen 3.8 27B Vision model (ranging from 8bit to 3.23bpw) and benchmarked them against community versions like lm-studio and mlx-community. Tested on an H200 with a custom dataset, their 11.8 GB DWQ quant achieved 70.32% top-1 agreement, significantly outperforming others of similar size. This version runs on a 16GB MacBook with a wired limit tweak, utilizing distillation from a BF16 teacher model.
More from Infra
- Open Source Facefusion Android App Runs Video Face Swap on Snapdragon NPU — Few_Caregiver8134 · 2026-08-26
- Apple M5 Mac Studio page features LM Studio for local AI — mattturck · 2026-08-26
- Chinese chip packaging firms invest $2.2B in expansion fueled by AI boom — pstAsiatech · 2026-08-26
- Can I run MiniMax H3 locally on an RTX 2060 with 6GB VRAM? — Amjad_K · 2026-08-26
- a16z Partner: Distinction between training and inference will become meaningless — Kyrannio · 2026-08-26
- Cerebras CEO explains the 3 supply-chain bottlenecks gripping AI chips: HBM, CoWoS, TSMC 3nm — rohanpaul_ai · 2026-08-26