Qwen3.8-Flash Runs Locally: 125B Model on Just 75GB RAM
danielhanchen · x · 2026-08-26
Unsloth announced that Qwen3.8-Flash is now available to run locally. This is a 125B parameter multimodal MoE model and an early preview of the Qwen4 architecture, outperforming Claude-4.6-Opus (Max).
Key Highlights:
- Low Resource Requirement: Runs on just 75GB RAM or unified memory with no GPU VRAM needed. The 1-bit quantized version is 79% smaller than BF16 while retaining top-1% accuracy.
- Architecture Upgrades: Features GDN + QSA hybrid attention, N-gram Embedding, and the Muon optimizer, drastically reducing inference costs.
- Tooling Support: Available via Unsloth GGUFs and a specific llama.cpp PR, making it highly suitable for Macs and devices with large memory capacities.
More from Infra
- Open-source Gepard TTS beats 24 closed APIs with 68ms latency on RTX 4090 — ylankgz · 2026-08-27
- Proposal: DGX Spark-class devices sold on $200/mo contracts could mesh into a giant cheap inference network — jasonkneen · 2026-08-27
- Glean reveals model routing scores: GPT-5.6 Luna leads at $0.08 — testingcatalog · 2026-08-27
- Serving frontier models at scale on purely Chinese hardware — tokumin · 2026-08-27
- Firecrawl Launches Startup Deal: Up to $30k in Credits — devdigest · 2026-08-27
- Long Read: AI Is Buying the Data of Dead Companies — rvp · 2026-08-27