Nativ local inference hits 11k tok/s prefill with LFM2.5 on M5 Max
helloiamleonie · x · 2026-08-05
Nativ, an open-source macOS app for running AI models locally on Apple Silicon, now supports Liquid AI's LFM2.5-2.6B in v0.2.2. On an M5 Max (48GB), it achieves 11,231 tok/s prefill and 84 tok/s decode in full bf16, with 128K context in 8.5GB. Batch 16 aggregate decode reaches 476 tok/s. The app offers chat, telemetry, multimodal support, and integrates with coding agents via MLX-VLM.
Related event: LiquidAI LFM2.5 Runs Locally on Mac with Strong Performance(3 posts)→
More from Apps
- Google forcibly replaces Assistant with Gemini, users report slow responses and lost features — Pyromethious · 2026-09-22
- Drowning in AI subscriptions, yet Google's AI search mode wins the usage battle — SuB8u · 2026-09-22
- Notion engineering lead confirms native iOS rewrite with more perf work coming — nbaschez · 2026-09-22
- Amazon blocks Meta's Muse AI agent from Amazon.com after Meta refuses removal — coolbern · 2026-09-22
- Stanford's VirtualBiotech puts tens of thousands of AI scientist agents in Science, NYT reports — StanfordAILab · 2026-09-22
- Early Muse hands-on: 'extremely underwhelming and useless', says econoar — econoar · 2026-09-22