New Qwen 3.8 MLX Engine Features Dynamic Quantization
MgkMshrmBrkfst · x · 2026-09-01
The HamsterResearch team is building a custom MLX engine for Qwen 3.8-Flash-Next. Key features include: downloading a single bf16 tiered model that can be quantized on-the-fly to bf16/q8/q4/q3; adjustable n-gram streaming from NVMe; and tunable expert residency ratios. The engine targets 130-150 tok/s aggregate decode speed to support efficient sub-agents.
Related event: Qwen3.8-Flash-Next with pMLX Engine Runs Big Models on Small Memory(3 posts)→
More from Infra
- SpaceX data center team shakeup: Musk replaces leaders with rocket, satellite internet execs — kyliebytes · 2026-09-02
- Seeking Datacenter-Grade OCS — jwt0625 · 2026-09-02
- GB10 price hike sparks debate: Is Mac Studio the best value for compute? — geekender · 2026-09-02
- Recurrent depth technique reduces memory and bandwidth costs — pstAsiatech · 2026-09-02
- Vercel adds AWS PrivateLink support for secure private connectivity — cramforce · 2026-09-02
- Hardware Advice: Running H3 Mini Max Locally on Mac — KourriX · 2026-09-02