Qwen3.8 Flash Next protip: use mmap mode for lazy loading
Pyrolistical · reddit · 2026-08-30
Sharing experience running Qwen3.8-Flash-Next in a memory-constrained environment. The default load-mode auto might not use mmap, leading to OOM. After manually enabling tensor-read-lazy on and load-mode mmap, the author successfully ran the Q4 quantized model on AMD GPUs and provided specific llama-bench commands and performance data (279 t/s).
More from Infra
- Qwen 350K Context Tested on M5 Max: Performance and Quality — Artistic_Okra7288 · 2026-08-30
- Azure Linux 4.0 Desktop Concept: PowerShell, Edge, and Copilot Pre-installed — unixterminal · 2026-08-30
- Jensen Huang: Built GPU tech first, found endless problems from graphics to molecular dynamics — r0ck3t23 · 2026-08-30
- How to build an LLM inference engine from scratch: 5-layer architecture — glenbeer · 2026-08-30
- Huaqin expects super node revenue to exceed 10B RMB in 2H 2026 — zephyr_z9 · 2026-08-30
- Nvidia is generating $1 billion a day, a business scale deemed absurd years ago — shauntrennery · 2026-08-30