Nativ local inference hits 11k tok/s prefill with LFM2.5 on M5 Max

helloiamleonie · x · 2026-08-05

Nativ, an open-source macOS app for running AI models locally on Apple Silicon, now supports Liquid AI's LFM2.5-2.6B in v0.2.2. On an M5 Max (48GB), it achieves 11,231 tok/s prefill and 84 tok/s decode in full bf16, with 128K context in 8.5GB. Batch 16 aggregate decode reaches 476 tok/s. The app offers chat, telemetry, multimodal support, and integrates with coding agents via MLX-VLM.

Related event: LiquidAI Releases LFM2.5 for High-Speed On-Device Mac Inference(2 posts)→

Original post →

More from Apps

Apps channel →