Testing Liquid AI's LFM2.5 2.6B: 90t/s Speed but Fails at Tool Calling
curiousily_ · reddit · 2026-08-06
A developer tested Liquid AI's LFM2.5 (2.6B dense model) using llama.cpp on an M5 Pro. Under Q8 quantization, the model achieved impressive speeds of 90 tokens/s and consumed only about 4GB of memory, highlighting its efficiency for local deployment.
However, the model underperformed significantly in agentic tasks. Compared to a competitor from its official benchmarks (Qwen3.5 4B at Q4), LFM2.5 struggled notably with tool calling. In the OpenCode environment, it had severe issues making actual tool calls and was genuinely confused about the working directory.
More from coding & agent
- A Vision for Cargo: Rust's Package Manager Eyes Agentic Development — charliermarsh · 2026-08-06
- Open-Source Benchmarks: RTX 5090 LLM Quants and 8GB VRAM Agentic Scores — max_paperclips · 2026-08-06
- Agent Harnesses Swing Accuracy by 15%: DataSpace Benchmark Reveals Shortcomings — omarsar0 · 2026-08-06
- Meta Launches Muse Code Terminal Coding Agent and Muse Spark 1.2 Model — alexandr_wang · 2026-08-06
- Christoph Nakazawa's Retrospective: Building a Framework with LLMs — cnakazawa · 2026-08-06
- Scale Details Muse Code: Co-training Model and Agent for Better Tool Use — alexandr_wang · 2026-08-06