M3 Ultra Hits 374 tok/s Decoding Running LiquidAI 3B Locally at 16 Concurrent Requests
JosephJacks_ · x · 2026-08-13
Nativ, an open-source local model runner, demonstrated the concurrent performance of running LiquidAI's LFM2.5-VL-3B multimodal model on Apple Silicon (M3 Ultra).
- High-Concurrency Decoding: Achieved 374.6 tok/s aggregate decode at 16 concurrent requests, which is 4.16× faster than a single request.
- Fast Prefill: Reached 5,494 tok/s prefill, 3.06× that of a single request.
- Fully Local: No cloud or API bills required; all data stays on the Mac. Available now in Nativ v0.3.1.
More from Infra
- OCP Debate on AI Compute Networking: Will Copper Survive or Will Optics Dominate? — BenBajarin · 2026-08-13
- Silicon Photonics Breakthrough: Chip Pbit Density Reaches 269,000 — beffjezos · 2026-08-13
- SpaceX Alumni Found Ambrosia Energy, Using Solar and Storage to Power AI Datacenters — Scobleizer · 2026-08-13
- Breaking the Memory Wall: How CXL Will Reshape AI Compute Infrastructure — BenBajarin · 2026-08-13
- Pure Rust Browser Port: Qwen3-TTS Runs Locally Without GPU — doodlestein · 2026-08-13
- Anthropic Reportedly in Talks to Acquire AI Chip Efficiency Startup Decart for ~$6B — coinfanking · 2026-08-13