MoEspresso runs Qwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max

marcobaldo · reddit · 2026-09-28

The author of MoEspresso, a hobby inference engine, released its third iteration: it runs the 125B Qwen3.8-Flash-Next locally on a 2021 32GB M1 Max at 12-15 decode tokens per second.

Key techniques:

On a reproducible 48-question benchmark, local Qwen3.8 Flash scored 84.3%, beating GPT-6 Luna xhigh (81.4) and Claude Opus 4.8 xhigh (80.3), trailing hosted Qwen3.8 Flash (89.7) and GPT-6 Sol medium (89.2).

Install via Homebrew; code is open source, with Linux and AMD Strix Halo support planned.

Original post →

More from Infra

Infra channel →