Liquid AI Releases 2.6B Model: 30 tok/s on Phone with 128K Context
BTA_Labs · reddit · 2026-08-05
Liquid AI released LFM2.5-2.6B, a small model designed for local devices. It features 2.69B parameters, 128K context, tool calling, and is post-trained for multi-step agent workflows.
Performance (Vendor Benchmarks):
- 30 tok/s on a phone, 220 tok/s on an M5 Max.
- Under 2.5 GB memory usage; Q4 GGUF is 1.67 GB.
Capabilities:
- Competitive in tool calling, but still lags behind larger models like Qwen3.5-9B in coding and knowledge-heavy tasks.
- The model card advises against using it for agentic coding.
The author suggests this model is ideal as a local worker agent for repetitive tasks like extraction and tool calls, leaving complex planning to larger models.
Related event: Liquid AI Launches LFM2.5-2.6B On-Device Agent Model(12 posts)→
More from Infra
- Reverse-Engineering NVIDIA Blackwell Tensor Cores for Bit-for-Bit Software Simulation — ycombinator · 2026-08-05
- The Bottleneck of AI Coding Isn't the Model, It's Your CI Pipeline — hichaelmart · 2026-08-05
- xAI Announces Fourth Data Center with 220,000 GB300 GPUs — chrisgrayson · 2026-08-05
- The Honest Cost Math: Moving Local LLM Stack to a Persistent Auto-Pausing GPU Desktop — ievseev · 2026-08-05
- Graph Analytics Benchmark Graph500 Selected for SPEC CPU 2026 Suite — Prof_DavidBader · 2026-08-05
- Groq Unveils LPX Architecture: Combines Its LPU with NVIDIA GPUs for Inference — GroqInc · 2026-08-05