llama.cpp PR Halves Indexer Score Memory for Qwen Flash, Cutting VRAM Use
jacek2023 · reddit · 2026-10-03
A merged pull request (qwen4exp) in ggml-org/llama.cpp halves the indexer score memory footprint, so Qwen Flash Next now runs with noticeably less VRAM. A direct win for users running Qwen locally on memory-constrained hardware.
More from Infra
- Fake Intel N150 mini PC scam exposed: seller hardcoded 'New_N150' into BIOS string — Ok-Shower7286 · 2026-10-03
- Crusoe CEO Chase Lochmiller on 20VC: $6.4B raised, GPU depreciation, and datacentre economics — 20VC · 2026-10-03
- Qwen 177B at 11-15 tok/s on a Single RTX 5070 12GB: Expert Streaming Deep Dive — ayobluestarr · 2026-10-03
- Cloudflare Durable Objects now survive client disconnects for long-running agents — threepointone · 2026-10-03
- SemiAnalysis: Nvidia's custom NVHBM frees ~25% more compute die area on Feynman — zephyr_z9 · 2026-10-03
- gufo-Qwen3.6-35B hits 3095 tok/s prefill, 190 tok/s decode on Strix Halo — nubela · 2026-10-03