Running LFM2.5-2.6B on OnePlus 13 Pure CPU at 17 tok/s

trikboomie · reddit · 2026-08-05

The author successfully ran the LFM2.5-2.6B model—a 2.69B parameter model with a 128K context window purpose-built for multi-step agent workflows—on a OnePlus 13 using purely the CPU.

Running the Q4KM GGUF version on a custom-built inference engine, the setup achieved a generation speed of about 17 tokens/s. The entire engine is only 450KB in size and supports other architectures like Qwen and Gemma. The author is currently aiming to push the performance to 30 tokens/s.

Original post →

More from Infra

Infra channel →