Running LFM2.5-2.6B on OnePlus 13 Pure CPU at 17 tok/s
trikboomie · reddit · 2026-08-05
The author successfully ran the LFM2.5-2.6B model—a 2.69B parameter model with a 128K context window purpose-built for multi-step agent workflows—on a OnePlus 13 using purely the CPU.
Running the Q4KM GGUF version on a custom-built inference engine, the setup achieved a generation speed of about 17 tokens/s. The entire engine is only 450KB in size and supports other architectures like Qwen and Gemma. The author is currently aiming to push the performance to 30 tokens/s.
More from Infra
- Anthropic to Develop Custom AI Chips for Faster and More Efficient Claude — Polymarket · 2026-08-05
- Wall Street Expects AI Capex Surge: OpenAI Quarterly Spending Could Top $18B — PTrubey · 2026-08-05
- Benchmarking MiniMax H3 on a 4090: Sage Attention Slashes Generation Time — thegr8anand · 2026-08-05
- 3090 Upgrade Dilemma: Is 24GB VRAM Enough or Jump to 32GB? — gtech02 · 2026-08-05
- Stanford Hazy Research: AI Agents Are Retiring CUDA Abstraction Layers — sumitdotml · 2026-08-05
- AI Agent Autonomously Writes Triton Kernel, Breaking NanoGPT Speedrun Record — cong_ml · 2026-08-05