Qwen3.8-Flash Runs Locally: 125B Model on Just 75GB RAM

danielhanchen · x · 2026-08-26

Unsloth announced that Qwen3.8-Flash is now available to run locally. This is a 125B parameter multimodal MoE model and an early preview of the Qwen4 architecture, outperforming Claude-4.6-Opus (Max).

Key Highlights:

Original post →

More from Infra

Infra channel →