Laguna XS 2.1 hits 200 decode TPS on M5 Max, up from 75 TPS
gajesh · x · 2026-08-12
Laguna XS 2.1 has achieved a decode speed of over 200 TPS on an M5 Max, a massive leap from just 75 TPS two weeks ago.
This breakthrough was driven by open-source community collaboration and promoted by @poolsideai. The team initially aimed for only 100 TPS, and this result sets a new performance benchmark for efficient on-device LLM inference.
More from Infra
- RTX 4080 hits OOM running Minimax Ref2V: How to generate long videos locally? — witcherknight · 2026-08-12
- Nvidia raises RTX 6000 Pro price to $16,000 on official website — Norwood_Reaper_ · 2026-08-12
- Temasek Plans Direct Investment in Samsung and SK Hynix, Chip Stocks Surge — firstadopter · 2026-08-12
- Musk on AI Compute Vision: Aiming for 10GW Next Year, Inference Heading to Space — elonmusk · 2026-08-12
- AI Data Center Load Growth Yields $5B in Savings for Texas Ratepayers — toptickcrypto · 2026-08-12
- Local MoE Benchmark: NVIDIA Lightning Outruns Qwen by 2.5x — parepeg · 2026-08-12