AtomicChat's Qwen3.8-Flash-Next Quant Cuts RAM from 106GB to 65GB, Prefill at 500 t/s
tolitius · reddit · 2026-08-29
A Reddit user tests AtomicChat's GGUF quant of Qwen3.8-Flash-Next on M4 Max 128GB, reducing memory from 106GB to 65GB with cold prefill 500 t/s. The quant uses llama.cpp mmap and pageable PLE. Also mentions oMLX PRs improving PLE SSD offload, nearly 3x faster cold prefill.
More from Infra
- DeepSeek V4 pricing outruns GLM-5.3 and Qwen-3.8 in cost-performance debate — teortaxesTex · 2026-08-29
- Poor EDA tool file compression creates SaaS opportunity — ai · 2026-08-29
- Unions support data centers, highlighting benefits over dismissal — apples_jimmy · 2026-08-29
- Qwen 3.8 Flash offloads ngram table to SSD via SGLang with zero loss — Easy_Werewolf7903 · 2026-08-29
- First vLLM Conference wraps with a capacity rooftop happy hour co-hosted by AMD — vllm_project · 2026-08-29
- Cerebras founder: AI is accelerating hardware evolution and reshaping the industry — Sethwinterroth · 2026-08-29