Qwen3.8 Variants: 8-bit Model Cuts Memory by 27%
EyalToledano · x · 2026-08-28
Two MLX quantized variants based on Qwen3.8-Flash-Next-REAP have been released.
- Qwen3.8-Flash-Next-REAP-288-MLX-8bit: 8-bit, 90.9% HumanEval, 70 GB resident (27% smaller than stock q4).
- Qwen3.8-Flash-Next-REAP-384-MLX-8bit: 4-bit with 384 experts, 92.1% HumanEval, 51 GB resident.
More from Infra
- Puro-2B Matches Qwen2.5 Performance with $6.9K Pretraining on RTX 5090 — _reachsumit · 2026-08-28
- Intel XE3P projected specs: 1.3 PFLOPS FP8, 1.5TB/s bandwidth, 2027 launch — QuixiAI · 2026-08-28
- Rumor: Anthropic interested in developing its own training chip — zephyr_z9 · 2026-08-28
- India Commits $13.4B for 'Semicon 2.0' Chip Design and Manufacturing — SumitGup · 2026-08-28
- Meta, Google, NTT to discuss AI data center optical architectures — jwt0625 · 2026-08-28
- Coherent-lite Catches Up to IMDD in Energy Efficiency for Pluggables — jwt0625 · 2026-08-28