Testing Liquid AI's LFM2.5 2.6B: 90t/s Speed but Fails at Tool Calling

curiousily_ · reddit · 2026-08-06

A developer tested Liquid AI's LFM2.5 (2.6B dense model) using llama.cpp on an M5 Pro. Under Q8 quantization, the model achieved impressive speeds of 90 tokens/s and consumed only about 4GB of memory, highlighting its efficiency for local deployment.

However, the model underperformed significantly in agentic tasks. Compared to a competitor from its official benchmarks (Qwen3.5 4B at Q4), LFM2.5 struggled notably with tool calling. In the OpenCode environment, it had severe issues making actual tool calls and was genuinely confused about the working directory.

Original post →

More from coding & agent

coding & agent channel →