llama.cpp PR adds --n-cpu-ffn option for dense model offload

jacek2023 · reddit · 2026-08-27

A new Pull Request in llama.cpp introduces the --n-cpu-ffn option. Similar to the existing --n-cpu-moe, this feature allows users to specify the number of FFN sublayers in dense models that should be offloaded to the CPU.

Original post →

More from Infra

Infra channel →