Laguna XS 2.1 hits 200 decode TPS on M5 Max, up from 75 TPS

gajesh · x · 2026-08-12

Laguna XS 2.1 has achieved a decode speed of over 200 TPS on an M5 Max, a massive leap from just 75 TPS two weeks ago.

This breakthrough was driven by open-source community collaboration and promoted by @poolsideai. The team initially aimed for only 100 TPS, and this result sets a new performance benchmark for efficient on-device LLM inference.

Original post →

More from Infra

Infra channel →