Nativ local inference hits 11k tok/s prefill with LFM2.5 on M5 Max
helloiamleonie · x · 2026-08-05
Nativ, an open-source macOS app for running AI models locally on Apple Silicon, now supports Liquid AI's LFM2.5-2.6B in v0.2.2. On an M5 Max (48GB), it achieves 11,231 tok/s prefill and 84 tok/s decode in full bf16, with 128K context in 8.5GB. Batch 16 aggregate decode reaches 476 tok/s. The app offers chat, telemetry, multimodal support, and integrates with coding agents via MLX-VLM.
Related event: LiquidAI Releases LFM2.5 for High-Speed On-Device Mac Inference(2 posts)→
More from Apps
- Voice AI Agent Evals and the Consumer Market Opportunity — agihouse_org · 2026-08-05
- Developer Builds Custom Music VST Plugins Using OpenAI Codex — Yamapama · 2026-08-05
- AI Tool Sparks Debate by Auto-Beautifying GitHub Contribution Graphs — aniketmaurya · 2026-08-05
- OpenAI Shares How o3 Deep Research Helps Solve Rare Pediatric Diseases — OpenAI · 2026-08-05
- Chrome Extension Uses WebGPU and Gemma for On-Device Text Summarization — jason_mayes · 2026-08-05
- OpenAI Engineer Details WebRTC Improvements Powering GPT-Live — juberti · 2026-08-05