llama.cpp v0.4.0 ships Qwen3.8-Flash-Next support, sparse flash attention, and RDMA
solyarisoftware · x · 2026-09-05
llama.cpp released v0.4.0 with notable additions:
- Initial architecture support for Qwen3.8-Flash-Next (qwen4exp) and NVIDIA Nemotron-3-Puzzle-75B-A9B
- New llamalazymode for on-demand tensor reading
- Per-slot server context limits
- Multimodal video input options and new mtmd APIs
- ggml bumped to 0.23.0 with major sparse flash attention and RDMA work
Also includes session/state version bumps and a maxbufsize quantization parameter.
More from Infra
- Jensen Huang: 1GW of AI data center costs $50-60B, and 'we're building 100 gigawatts' by decade's end — victor_explore · 2026-09-05
- Full recipe: running Qwen3.8 27B on AMD Strix Halo with patched ROCm llama.cpp — ilintar · 2026-09-05
- Open-sourced Lightpanda: headless browser 11x faster than Chrome with 9x less RAM — JafarNajafov · 2026-09-05
- Tencent Hunyuan preview: 770B params, 1M context, Apache 2.0, 214GiB quantized — Aiden_Tech_Ai · 2026-09-05
- NVIDIA's $1B Single-Model Training Forecast Landed Years Ahead of Schedule — IgorCarron · 2026-09-05
- Starting local AI on attic hardware: 3x NUC11 plus Ryzen 3700X rig — -markusb- · 2026-09-05