Qwen3.8-Flash-Next Architecture: Estimated Local Memory ~82GB

pmv143 · reddit · 2026-08-26

A Reddit user analyzes the unreleased Qwen3.8-Flash-Next architecture (125B-A6B + 51B n-gram). Ideal 4-bit quantization requires 82GB VRAM (58GB main weights + 24GB n-gram tables). Due to sparse access, the n-gram table is an excellent candidate for system RAM offload, potentially making it very local-friendly upon release.

Original post →

More from Infra

Infra channel →