LiquidAI's LFM2.5-2.6B Hits 82 tok/s Decode on Mac with 128K Context
helloiamleonie · x · 2026-08-05
LiquidAI released LFM2.5-2.6B, a model built for agentic and coding workflows, fully optimized for on-device execution with 128K context.\n\nBenchmarked locally on an M5 Max (48GB) via Nativ using full bf16 with no quantization, the model delivers impressive metrics: 11,231 tok/s prefill and 82 tok/s decode. It fits the full 128K context in just 8.5GB of memory, achieving 476 tok/s aggregate decode at batch 16.
More from Infra
- AMD Claims Helios Rack Beats Nvidia's Vera Rubin by 30% in Tokens per Dollar — Beth_Kindig · 2026-08-05
- Burning $130K/Day? Unpacking DeepSeek API Token Volumes — teortaxesTex · 2026-08-05
- NSF Launches $100M Program for Regional AI Infrastructure Hubs — mkratsios47 · 2026-08-05
- Chutes AI Enforces TEE Verification: 8x RTX 5090s Beat Pro GPUs at 65% Lower Cost — markjeffrey · 2026-08-05
- engyai Launches Cheapest Kimi K3 API on OpenRouter, Cutting Costs by 50% — const_reborn · 2026-08-05
- ai& Partners with Voltaiq for Battery Storage in Japanese AI Data Centers — DavidBennett__ · 2026-08-05