Moonshot’s Kimi K3 goes live on Phala with 2.8T parameters and 1M-token context
bgmshana · x · 2026-07-29
Phala says Moonshot AI’s Kimi K3 is now live on its confidential inference stack.
The model is presented as a 2.8T-parameter open-weight multimodal reasoning model with a 1M-token context window, aimed at complex coding, knowledge work, and long-horizon agent workflows. Phala also lists GPU TEE-backed inference, pricing of $3/M input and $15/M output, and a text+image input shape with text output.
The page also compares the private inference route against alternatives and highlights the confidential-computing angle as part of the deployment story.
More from Infra
- YouTube-style semantic IDs tackle recommender memory walls with dual-purpose tokens — _reachsumit · 2026-07-29
- Meta’s Memory Layer pushes Instagram Reels item coverage to 100% and freshness to 20 seconds — _reachsumit · 2026-07-29
- VaLiDRec uses variable-length LLM-aligned IDs and runs 87.49× faster than LC-Rec — _reachsumit · 2026-07-29
- India’s AI inference market will reward companies that co-optimize models and hardware — santoshpanda · 2026-07-29
- Hermes Agent Desktop impresses users with parallel tools and remote local-model setup — Teknium · 2026-07-29
- Chinese threat actor pivots infrastructure and leaks 775 API-key IDs from an AI reseller — cyb3rops · 2026-07-29