AtomicChat's Qwen3.8-Flash-Next Quant Cuts RAM from 106GB to 65GB, Prefill at 500 t/s

tolitius · reddit · 2026-08-29

A Reddit user tests AtomicChat's GGUF quant of Qwen3.8-Flash-Next on M4 Max 128GB, reducing memory from 106GB to 65GB with cold prefill 500 t/s. The quant uses llama.cpp mmap and pageable PLE. Also mentions oMLX PRs improving PLE SSD offload, nearly 3x faster cold prefill.

Original post →

More from Infra

Infra channel →