New Qwen 3.8 MLX Engine Features Dynamic Quantization

MgkMshrmBrkfst · x · 2026-09-01

The HamsterResearch team is building a custom MLX engine for Qwen 3.8-Flash-Next. Key features include: downloading a single bf16 tiered model that can be quantized on-the-fly to bf16/q8/q4/q3; adjustable n-gram streaming from NVMe; and tunable expert residency ratios. The engine targets 130-150 tok/s aggregate decode speed to support efficient sub-agents.

Related event: Qwen3.8-Flash-Next with pMLX Engine Runs Big Models on Small Memory(3 posts)→

Original post →

More from Infra

Infra channel →