llama.cpp adds tensor split for LFM2/MoE, boosting inference performance significantly

pmttyji · reddit · 2026-08-21

A PR submitted to the llama.cpp repository adds --split-mode tensor support for the LFM2 and LFM2MoE model families, accompanied by detailed benchmarks on 2x RTX A5000 (NVLink).

Feature Update:

Accuracy & Performance Benchmarks:

Original post →

More from Infra

Infra channel →