Qwen3.8-Flash-Next runs 1M context locally on M5 Max at 40 tok/s via MLX-serve

Beamsters · reddit · 2026-09-09

The co-creator of the new Qwen3.8-Flash-Next engine support in MLX-serve shares a local deployment achieving 1M-token context on an M5 Max 128GB:

All resources open-sourced: the mlx-serve engine, model weights on HuggingFace, and the Opencode2 plugin. Bugs expected — reports welcome.

Original post →

More from Infra

Infra channel →