Third-Party Devs Boost Kimi K3 Inference to Nearly 80 TPS, Beating Official API

bittingthembits · x · 2026-07-30

According to a leak by TheTNetHunter, a dev team @engyai is set to release an optimized Kimi K3 inference service within 48 hours. They have successfully pushed the model's generation speed from an initial 20 TPS up to 50-78 TPS.

This performance not only surpasses the official Kimi API but also matches inference efficiencies comparable to running on B300 GPUs, offering the market a highly cost-effective alternative for Kimi K3 inference.

Original post →

More from Infra

Infra channel →