Running Qwen3.8-Next-Flash on 96GB RAM: offload the n-gram table to SSD
Iory1998 · reddit · 2026-08-30
A user with dual 5070 Ti + 3090 (40GB VRAM) and 96GB RAM shares a hands-on debugging story running Qwen3.8-Next-Flash Q4 locally:
- Unsloth quants failed: Q4KXL (105GB) and IQ4XS (90GB) bake the 56B n-gram table into the weight files, so everything loads into memory; with KV cache, conversations beyond 60K–100K tokens crash llama.cpp silently, at only 12–14 t/s.
- The AtomicChat build is the key difference: it shards the n-gram table onto SSD, needing only 54GB of fast memory. The table is touched sparsely (2.7KB per token via hash), so SSD offload costs almost nothing.
- Results: a real 217K-token conversation ran end-to-end without crashing, full 256K context fits, and speed rose to 22 t/s. Prefill remains hardware-bound at 18 min for 220K tokens.
- Fun fact: the AtomicChat option was suggested by DeepSeek running benchmarks on his own rig. The author cautions quant quality parity with unsloth isn't guaranteed.
More from Infra
- Oracle nears $1T market cap as AI infrastructure captures top value — thedealdirector · 2026-08-30
- Grok Bot Provides Cloud Desktop Environment for Bots — tristanbob · 2026-08-30
- Token Economics: How Cache & Latency Impact Pricing — AccBalanced · 2026-08-30
- NVIDIA Studio Driver Optimizes ComfyUI and LTX-2.5 Performance — PixWizardry · 2026-08-30
- DLSS 5 Video Player tested: Runs slow on 3090 Ti due to FP8 lack — fallengt · 2026-08-30
- Data Centers Drive US Reindustrialization: Benefits from Taxes to Jobs — GavinSBaker · 2026-08-30