LiquidAI's LFM2.5-2.6B Hits 82 tok/s Decode on Mac with 128K Context

helloiamleonie · x · 2026-08-05

LiquidAI released LFM2.5-2.6B, a model built for agentic and coding workflows, fully optimized for on-device execution with 128K context.\n\nBenchmarked locally on an M5 Max (48GB) via Nativ using full bf16 with no quantization, the model delivers impressive metrics: 11,231 tok/s prefill and 82 tok/s decode. It fits the full 128K context in just 8.5GB of memory, achieving 476 tok/s aggregate decode at batch 16.

Original post →

More from Infra

Infra channel →