LiquidAI's LFM2.5-2.6B Hits 82 tok/s Decode on Mac with 128K Context
helloiamleonie · x · 2026-08-05
LiquidAI released LFM2.5-2.6B, a model built for agentic and coding workflows, fully optimized for on-device execution with 128K context.\n\nBenchmarked locally on an M5 Max (48GB) via Nativ using full bf16 with no quantization, the model delivers impressive metrics: 11,231 tok/s prefill and 82 tok/s decode. It fits the full 128K context in just 8.5GB of memory, achieving 476 tok/s aggregate decode at batch 16.
Related event: LiquidAI LFM2.5 Runs Locally on Mac with Strong Performance(3 posts)→
More from Infra
- StepFun's Step 5 Preview scores 44 on AA Intelligence Index at ~2.8x lower cost than Kimi K3 — ArtificialAnlys · 2026-09-22
- SemiAnalysis tears down A20 on TSMC N2, DRAM revealed in iPhone 18 Pro Max package — dylan522p · 2026-09-22
- Bernstein's Intel road notes: servers sold out through 2027, real test comes in 2028 — BenBajarin · 2026-09-22
- Rollout scheduling: the underappreciated infra trick boosting inference efficiency — stochasticchasm · 2026-09-22
- Grok reportedly hit by a datacenter incident, details still unclear — Daniel_Farinax · 2026-09-22
- Decoupled agent runtime, mini harnesses: notes on a frontier lab's infra stack — stochasticchasm · 2026-09-22