llama.cpp PR adds --n-cpu-ffn option for dense model offload
jacek2023 · reddit · 2026-08-27
A new Pull Request in llama.cpp introduces the --n-cpu-ffn option. Similar to the existing --n-cpu-moe, this feature allows users to specify the number of FFN sublayers in dense models that should be offloaded to the CPU.
More from Infra
- Nvidia Reportedly Agrees to Buy Hugging Face for $12.9 Billion — ferruz_noelia · 2026-08-27
- AI generates 500k words on 31 kWh, equaling 310 human hours of energy — cis_female · 2026-08-27
- Cursor User Burns 472.8B Tokens in a Month, Highlighting Cost Limits — Daniel_Farinax · 2026-08-27
- Can 96GB Mac Studio run Qwen3.8? Analyzing SSD offload feasibility — Mxmtm · 2026-08-27
- Deep Dive: AWS S3 Architecture, Rust Rewrite, and Heat Management at 280 Trillion Objects — Franc0Fernand0 · 2026-08-27
- Prefix Sliding: discarding stale reasoning tokens makes test-time scaling 3x faster — Bedrovelsen · 2026-08-27