Inco AI Launches Inference Platform, Leading Output Speed on Kimi K3, MiniMax M3, GLM 5.3
songhan_mit · x · 2026-09-04
Inco AI launched its inference platform purpose-built for the agentic era, offering high-speed endpoints for Kimi K3 (287 tok/s, 2.0× next-fastest), MiniMax M3 (417 tok/s, 1.7×), GLM 5.3 (401 tok/s, 1.3×), and GLM 5.3 Flash (458 tok/s, 1.4×)—each topping Artificial Analysis provider speed leaderboards. The stack combines fleet-scale GPU serving with cache-aware routing, in-house DFlash/DFlash 2 block-diffusion speculative decoding (open-source checkpoints available), and AWQ/ParoQuant quantization research.
More from Infra
- Starship fires all 33 Raptors at once, hitting 74M newtons—double Saturn V — PeterDiamandis · 2026-09-04
- xAI wins approval for fifth Memphis-area data center in $40M land swap with Southaven — chrisgrayson · 2026-09-04
- Mirai's uzu engine brings speculative decoding to Apple M5, hitting 105 tok/s on Qwen3.6 27B — TheMoonMidas · 2026-09-04
- Best local models for 12GB of VRAM: Gemma-4-12B remains the pick — GlennCameronjr · 2026-09-04
- Rumor: When One Major AI Service Goes Down, Its Traffic Takes the Rest Down With It — jxnlco · 2026-09-04
- UBS: Ant is now Broadcom's second-largest customer after Google; analyst bets OpenAI works more with MediaTek — BenBajarin · 2026-09-04