Homegrown inference engine runs Xiaomi's 1T-parameter model at 1300 tok/s
bookwormengr · x · 2026-09-30
In a reply, @bingxu claims their in-house generated inference engine can run Xiaomi's 1T-parameter model at 1300 tokens per second, with a demo video attached. If accurate, that would be a striking throughput for such a massive MoE model, though no implementation details or hardware specs are disclosed in the post.
More from Infra
- Kafgres embeds a Rust Kafka broker inside Postgres, clients can't tell the difference — JiliJeanlouis · 2026-09-30
- Starlink Mobile goes live in Ecuador, bringing direct-to-cell coverage to the Amazon and Galápagos — elonmusk · 2026-09-30
- OpenRouter data: token usage exploding, some open-weight models see 10x spend since January — AccBalanced · 2026-09-30
- Reserve unlocked GPU capacity for startups, not just big logos — AccBalanced · 2026-09-30
- Ollama 0.35 ships with local Nimble support on Mac — Technovangelist · 2026-09-30
- Nvidia CEO Jensen Huang: Data centers are now 'superintelligence factories' — Polymarket · 2026-09-30