LFM2.5 2.6B beats MiniCPM5 2B in speed and RAM: 22 t/s vs 16 t/s on M1 Air
parepeg · reddit · 2026-10-05
A Reddit user benchmarked LFM2.5 2.6B and MiniCPM5 2B on small agentic tool-use tasks (e.g. weather queries) running locally on an M1 Air.
Verdict: LFM2.5 wins
- Faster: 200 t/s prefill and 22 t/s generation vs MiniCPM5's 162/16 t/s
- Leaner: 2.5GB RAM at 32k context vs MiniCPM5's 3.8GB despite fewer parameters
- Good at one-shot agentic tasks, but hallucinates as conversations grow longer
MiniCPM5
- Ships with a DSpank draft model for speculative decoding
- Often replies in Chinese when prompted in English, yet seems smarter in English and potentially stronger at agentic work
The post includes full llama-server launch commands (quantized KV cache, draft model flags, sampling settings) so the setup is reproducible.
More from Infra
- TernaryQuench: open-source ternary quantization trainer for Qwen3 with MLX export — casper_hansen_ · 2026-10-05
- One USB-C cable turns an iPhone into a 24GB MacBook's extra memory, running local Qwen 27B 40%+ faster — TheMoonMidas · 2026-10-05
- Hugging Face Accelerate lead touts 6 years of inference work, v0.20.0 features — TheZachMueller · 2026-10-05
- Compute deals insider: buyers want only NVIDIA gear, CUDA moat alive and well — sudoraohacker · 2026-10-05
- Micron CEO: memory supply will be much tighter in 2027 and 2028 than in 2026 — chillinewman · 2026-10-05
- Indie dev's cheapest stack: Firebase analytics + Cloudflare R2 + Vercel — jdluk87 · 2026-10-05