Homegrown inference engine runs Xiaomi's 1T-parameter model at 1300 tok/s

bookwormengr · x · 2026-09-30

In a reply, @bingxu claims their in-house generated inference engine can run Xiaomi's 1T-parameter model at 1300 tokens per second, with a demo video attached. If accurate, that would be a striking throughput for such a massive MoE model, though no implementation details or hardware specs are disclosed in the post.

Related event: Two-Person Team Uses AI to Generate 20 Inference Engines in Two Weeks, Hitting 1300 tok/s(3 posts)→

Original post →

More from Infra

Infra channel →