llama.cpp PR adds CPU offload for dense model FFN layers
pmttyji · reddit · 2026-08-19
A new PR for llama.cpp proposes adding --n-cpu-ffn for dense models, similar to the existing MoE option. This offloads FFN layers to the CPU, enabling larger dense models (like Qwen 3.8 27B) to run on VRAM-constrained devices (e.g., 16GB). Benchmarks show running the 27B model at Q4KM with 130k context and 20 t/s.
More from Infra
- Marvell grants Google warrant as part of expanded custom AI chip deal — firstadopter · 2026-08-19
- Marvell signs custom chip deal with Google spanning the TPU ecosystem for AI inference — firstadopter · 2026-08-19
- SALT: CELF-Based Sentence-Level Compression for KV Cache Retrieval — No_Sky9786 · 2026-08-19
- Beyond human intuition: AI designs chip components 500x smaller than engineering limits — ChuckDBrooks · 2026-08-19
- ComfyUI becomes unusable overnight with MiniMax H3, causing system freezes — Fit-Association-448 · 2026-08-19
- Suggestion: Move Anthropic bio AI to Tenstorrent to cut costs — DavidBennett__ · 2026-08-19