llama.cpp switches lazy-mode default to auto: 51B-param embedding table stays on disk
whiteh4cker · reddit · 2026-09-01
Starting with b10726, llama.cpp's default --lazy-mode is now auto: the 51B-parameter PLE n-gram embedding table of Qwen 3.5 Flash Next is no longer loaded into RAM but mmap'd and read from disk on demand during inference, even with --load-mode none. Pass --lazy-mode off to restore the old eager loading behavior.
More from Infra
- Nvidia Uses Proxy Models Like 'GPTOSS 2T' to Benchmark for GPT-5 — nrehiew_ · 2026-09-01
- GPU Depreciation Paradox: Token Revenue vs. Resale Value — AccBalanced · 2026-09-01
- From $18K PC to Home Datacenter: The Escalating Compute Demand — kevinnbass · 2026-09-01
- Major Microsoft Outage Hits 365, Azure, Teams Globally — cyb3rops · 2026-09-01
- Data Center Controversies: Lack of Transparency and NDAs — cremieuxrecueil · 2026-09-01
- Debunking Anti-Data Center Talking Points: Water, Power, and Taxes — cremieuxrecueil · 2026-09-01