Inco AI Launches Inference Platform, Leading Output Speed on Kimi K3, MiniMax M3, GLM 5.3

songhan_mit · x · 2026-09-04

Inco AI launched its inference platform purpose-built for the agentic era, offering high-speed endpoints for Kimi K3 (287 tok/s, 2.0× next-fastest), MiniMax M3 (417 tok/s, 1.7×), GLM 5.3 (401 tok/s, 1.3×), and GLM 5.3 Flash (458 tok/s, 1.4×)—each topping Artificial Analysis provider speed leaderboards. The stack combines fleet-scale GPU serving with cache-aware routing, in-house DFlash/DFlash 2 block-diffusion speculative decoding (open-source checkpoints available), and AWQ/ParoQuant quantization research.

Original post →

More from Infra

Infra channel →