Moonshot Releases 2.8T-Parameter Kimi K3; Modal Achieves 460 TPS with Speculative Decoding
sarahcat21 · x · 2026-07-30
Moonshot released Kimi K3, its latest open-weight multimodal model, now available on Modal. K3 features 2.8 trillion parameters, a 1M token context window, and native vision, ranking as the top open model on public intelligence indexes.
Architecture & Performance
K3 uses a Mixture-of-Experts architecture with 16 experts active per token. It introduces Kimi Delta Attention and Attention Residuals, delivering roughly 2.5x the scaling efficiency of K2, specifically built for long-horizon agentic tasks.
Inference Optimization
Modal provided Day 0 vLLM support and custom-trained a DFlash speculator tuned to K3's architecture, achieving 460 tokens/s. Users can access it via the Shared API with token-based pricing.
Related event: Moonshot Releases 2.8T Open-Source Model Kimi K3(2 posts)→
More from Infra
- Sam Altman Understands Why People Don't Want AI Data Centers in Their Backyards — businessinsider · 2026-07-30
- Dual GPU inference with RTX 4090 + 3060: speed impact and optimization tips — cosmoschtroumpf · 2026-07-30
- Future 100T Param Model to Cost >$250B to Train, Says Joseph Jacks — JosephJacks_ · 2026-07-30
- Inference-Time Compute is the New Scaling Law Frontier, Says CoreWeave Exec — agihouse_org · 2026-07-30
- Benchmarking C++ vs PyTorch for RLHF Reward Model Inference — Venkata Naga Sai Vishnu Rohit Pulipaka · 2026-07-30
- Meta Projects Over $130 Billion in 2026 Capex, Stock Plunges 9% After-Hours — Polymarket · 2026-07-30